AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,737 results
20 May 2026

When Skills Don't Help: A Negative Result on Procedural Knowledge for Tool-Grounded Agents in Offensive Cybersecurity

AgentsDGX agent

arXiv:2605.20023v1 Announce Type: new Abstract: Agent Skills, structured packages of procedural knowledge loaded into an LLM agent at inference time, are widely reported to improve task pass rates by

WisdomAI’s new analytics agents go beyond insights, automating business work through autonomous action

AgentsDGX agent

WisdomAI Inc., creator of an artificial intelligence-native business intelligence platform, is jumping on the agentic AI bandwagon with the latest update to its flagship Federated Agentic Intelligence

19 May 2026

Before enterprises can run with agentic AI, they need to learn to walk with their data

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
AgentsDGX agent

Multi-agent orchestration is the destination, but for most enterprises, the road is blocked long before the first agent gets deployed by the quality of the data feeding those systems. As organizations

Can LLM Agents Be CFOs? Benchmarking Long-Horizon Resource Allocation in an Uncertain Enterprise Environment

AgentsDGX agent

arXiv:2603.23638v2 Announce Type: replace Abstract: Large language model (LLM) agents are increasingly tested on complex tasks, but their ability to allocate scarce resources over long horizons remain

GRASP: Graph Agentic Search over Propositions for Multi-hop Question Answering

AgentsDGX agent

arXiv:2605.16598v1 Announce Type: cross Abstract: Agentic retrieval improves multi-hop question answering by giving language models autonomy to iteratively gather evidence. Recent work augments these

Helpful to a Fault: Measuring Illicit Assistance in Multi-Turn, Multilingual LLM Agents

SafetyDGX agent

arXiv:2602.16346v3 Announce Type: replace Abstract: LLM-based agents execute real-world workflows via tools and memory. These affordances enable ill-intended adversaries to also use these agents to ca

MetaCogAgent: A Metacognitive Multi-Agent LLM Framework with Self-Aware Task Delegation

Model ReleasesDGX agent

arXiv:2605.17292v1 Announce Type: new Abstract: Multi-agent large language model (LLM) systems have shown promise for solving complex tasks through agent collaboration. However, existing frameworks as

NeuroMAS: Multi-Agent Systems as Neural Networks with Joint Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.16757v1 Announce Type: new Abstract: Multi-agent language systems are often built as hand-designed workflows, where agents are assigned semantic roles and communication protocols are specif

Stop rogue AI: How Unity Catalog secures your agent actions

AgentsDGX agent

Unity Catalog is a Databricks governance solution that helps secure AI agent actions by managing access controls and permissions across agent operations. It prevents unauthorized or 'rogue' AI behavio

Visual Agentic Memory: Enabling Online Long Video Understanding via Online Indexing, Hierarchical Memory, and Agentic Retrieval

Model ReleasesDGX agent

arXiv:2605.16481v1 Announce Type: cross Abstract: Long video understanding requires more than large context windows. It also needs a memory mechanism that decides what visual evidence to retain, keeps

18 May 2026

Differentiable Mixture-of-Agents Incentivizes Swarm Intelligence of Large Language Models

AgentsDGX agent

arXiv:2605.15706v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have catalyzed the development of multi-agent systems (MAS) for complex reasoning tasks. However, existi

Every time I ask my 10-year-old to use coding agents, he gets extremely disappointed. It turns out that all he wants is to build his own roc…

AgentsDGX agent

Every time I ask my 10-year-old to use coding agents, he gets extremely disappointed. It turns out that all he wants is to build his own rocket simulator. No amount of context engineering helps. No mo

H-Mem: A Novel Memory Mechanism for Evolving and Retrieving Agent Memory via a Hybrid Structure

AgentsDGX agent

arXiv:2605.15701v1 Announce Type: cross Abstract: Memory data are ubiquitous in Large Language Model (LLM)-based agents (e.g., OpenClaw and Manus). A few recent works have attempted to exploit agents'

Look Before You Leap: Autonomous Exploration for LLM Agents

AgentsDGX agent

arXiv:2605.16143v1 Announce Type: new Abstract: Large language model based agents often fail in unfamiliar environments due to premature exploitation: a tendency to act on prior knowledge before acqui

The Open Agent Leaderboard

AgentsDGX agent

The Open Agent Leaderboard is a benchmarking system hosted on Hugging Face that evaluates and ranks AI agents based on their performance across various tasks and capabilities. It provides a standardiz

when evaluating long running agents, all of your evals don't need to be end to end. i'm working on a proper blog about this, but in our eval…

AgentsDGX agent

when evaluating long running agents, all of your evals don't need to be end to end. i'm working on a proper blog about this, but in our evals for our agents that run for 30-60 minutes, we have two set

15 May 2026

Known By Their Actions: Fingerprinting LLM Browser Agents via UI Traces

AgentsDGX agent

arXiv:2605.14786v1 Announce Type: cross Abstract: As LLM-based agents increasingly browse the web on users' behalf, a natural question arises: can websites passively identify which underlying model po

SuperGrok now in Hermes Agent

AgentsDGX agent

SuperGrok has been integrated into the Hermes Agent, representing an advancement in Nous Research's AI agent capabilities. This integration likely combines SuperGrok's reasoning or processing features

Web Agents Should Adopt the Plan-Then-Execute Paradigm

Model ReleasesDGX agent

arXiv:2605.14290v1 Announce Type: cross Abstract: ReAct has become the default architecture across LLM agents, and many existing web agents follow this paradigm. We argue that it is the wrong default

You can now use your @grok subscription inside @NousResearch Hermes Agent. http://x.ai/news/grok-hermes

AgentsDGX agent

Nous Research has integrated Grok, xAI's large language model, into the Hermes Agent framework, allowing users with active @grok subscriptions to leverage Grok's capabilities within the Hermes Agent e

14 May 2026

AgenticFlict: A Large-Scale Dataset of Merge Conflicts in AI Coding Agent Pull Requests on GitHub

AgentsDGX agent

arXiv:2604.03551v2 Announce Type: replace-cross Abstract: Software Engineering 3.0 marks a paradigm shift in software development, in which AI coding agents are no longer just assistive tools but acti

CHAL: Council of Hierarchical Agentic Language

AgentsDGX agent

arXiv:2605.12718v1 Announce Type: new Abstract: Multi-agent debate has emerged as a promising approach for improving LLM reasoning on ground-truth tasks, yet current methodologies face certain structu

Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue

Model ReleasesDGX agent

arXiv:2605.12920v1 Announce Type: cross Abstract: Effective collaboration between embodied agents requires more than acting in a shared environment; it demands communication grounded in each agent's e

Excited to see SWE-ZERO trending on HF alongside awesome agentic trace datasets like AgentTrove!

AgentsDGX agent

SWE-ZERO and AgentTrove are trending datasets on Hugging Face related to software engineering and agentic AI systems. The post indicates growing interest in trace datasets that capture agent behaviors

Freshworks unveils Freddy AI Agent Studio and MCP Gateway for Freshservice

AgentsDGX agent

Freshworks Inc. today unveiled an expanded set of agentic capabilities in its Freshservice information technology service management platform led by a new no-code Freddy AI Agent Studio that lets ente

Introducing the Anyscale Agent Skill for LLM Post-Training

AgentsDGX agent

Anyscale has introduced a new Agent Skill designed to enhance LLM post-training capabilities, likely enabling developers to build and train agentic systems more effectively using Ray's distributed com

OpenAaaS: An Open Agent-as-a-Service Framework for Distributed Materials-Informatics Research

AgentsDGX agent

arXiv:2605.13618v1 Announce Type: cross Abstract: The Materials Genome Initiative catalyzed the proliferation of centralized platforms--SaaS, PaaS, and IaaS--that aggregate computational and experimen

The Horizon Threshold in Cooperative Multi-Agent Reward-Free Exploration

SafetyDGX agent

arXiv:2602.01453v3 Announce Type: replace Abstract: We study cooperative multi-agent reinforcement learning in the setting of reward-free exploration, where multiple agents jointly explore an unknown

13 May 2026

Focusing Influence Mechanism for Multi-Agent Reinforcement Learning

AgentsDGX agent

arXiv:2506.19417v2 Announce Type: replace Abstract: Cooperative multi-agent reinforcement learning (MARL) under sparse rewards remains fundamentally challenging because agents often fail to concentrat

just tried this out and it one-shotted* this video: 'before the agent does anything' *i generated the narrative using chatgpt and used that …

AgentsDGX agent

just tried this out and it one-shotted* this video: 'before the agent does anything' *i generated the narrative using chatgpt and used that as a prompt. featuring: @e2b @runanywhereai @composio @mem0a

12 May 2026

EvoMAS: Learning Execution-Time Workflows for Multi-Agent Systems

SafetyDGX agent

arXiv:2605.08769v1 Announce Type: new Abstract: Large language model (LLM)-based multi-agent systems have shown strong potential on complex tasks through agent specialization, tool use, and collaborat

GAMBIT: A Three-Mode Benchmark for Adversarial Robustness in Multi-Agent LLM Collectives

Model ReleasesDGX agent

arXiv:2605.09027v1 Announce Type: new Abstract: In multi-agent systems (MAS), a single deceptive agent can nullify all gains of an agentic AI collective and evade deployed defenses. However, existing

MIND-Skill: Quality-Guaranteed Skill Generation via Multi-Agent Induction and Deduction

AgentsDGX agent

arXiv:2605.08670v1 Announce Type: new Abstract: Large language model (LLM) powered AI agents have emerged as a promising paradigm for autonomous problem-solving, yet they continue to struggle with com

Robust Multi-Agent LLMs under Byzantine Faults

AgentsDGX agent

arXiv:2605.09076v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly collaborate over peer-to-peer networks to improve their reliability. However, these same interactions c

WebTrap: Stealthy Mid-Task Hijacking of Browser Agents During Navigation

AgentsDGX agent

arXiv:2605.08310v1 Announce Type: cross Abstract: Browser agents are increasingly deployed in long-horizon tasks, which require executing extended action chains to accomplish user goals. However, this

Willful Disobedience: Automatically Detecting Failures in Agentic Traces

AgentsDGX agent

arXiv:2603.23806v2 Announce Type: replace-cross Abstract: AI agents are increasingly embedded in real software systems, where they execute multi-step workflows through multi-turn dialogue, tool invoca

11 May 2026

SOM: Structured Opponent Modeling for LLM-based Agents via Structural Causal Model

AgentsDGX agent

arXiv:2605.07301v1 Announce Type: new Abstract: Accurately predicting opponents' behavior from interactions is a fundamental capability for large language model (LLM)-based agents in multi-agent and g

WebClipper: Efficient Evolution of Web Agents with Graph-based Trajectory Pruning

AgentsDGX agent

arXiv:2602.12852v2 Announce Type: replace Abstract: Deep Research systems based on web agents have shown strong potential in solving complex information-seeking tasks, yet their search efficiency rema

8 May 2026

Ship code within minutes with the Gemini CLI DevOps Extension

Model ReleasesDGX agent

With AI coding tools like Antigravity and Claude Code, I can build a working web app in record time. But deploying it? That's where I'd historically lose the rest of the afternoon to Dockerfiles, IAM

7 May 2026

Agentic publications: redesigning scientific publishing in the age of thinking large language models

AgentsDGX agent

arXiv:2505.13246v2 Announce Type: replace Abstract: Purpose: This paper introduces the concept of 'Agentic Publication,' a novel LLM-driven framework designed to complement traditional scientific publ

OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents

AgentsDGX agent

arXiv:2605.05185v1 Announce Type: new Abstract: Deep search has become a crucial capability for frontier multimodal agents, enabling models to solve complex questions through active search, evidence v

6 May 2026

A Low-Latency Fraud Detection Layer for Detecting Adversarial Interaction Patterns in LLM-Powered Agents

AgentsDGX agent

arXiv:2605.01143v1 Announce Type: new Abstract: Large Language Model (LLM)-powered agents demonstrate strong capabilities in autonomous task execution, tool use, and multi-step reasoning. However, the

Beyond State Machines: Executing Network Procedures with Agentic Tool-Calling Sequences

AgentsDGX agent

arXiv:2605.02584v1 Announce Type: cross Abstract: Agentic AI will be an essential enabling technology for designing future mobile communication systems, which could provide flexible and customized ser

Catching the Infection Before It Spreads: Foresight-Guided Defense in Multi-Agent Systems

SafetyDGX agent

arXiv:2605.01758v1 Announce Type: new Abstract: Large multimodal model-based Multi-Agent Systems (MASs) enable collaborative complex problem solving through specialized agents. However, MASs are vulne

CoFlow: Coordinated Few-Step Flow for Offline Multi-Agent Decision Making

AgentsDGX agent

arXiv:2605.01457v1 Announce Type: new Abstract: Generative models have emerged as a major paradigm for offline multi-agent reinforcement learning (MARL), but existing approaches require many iterative

Foundation-Model-Based Agents in Industrial Automation: Purposes, Capabilities, and Open Challenges

AgentsDGX agent

arXiv:2605.02592v1 Announce Type: new Abstract: Foundation models, particularly large language models, are increasingly integrated into agent architectures for industrial tasks such as decision suppor

Towards Multi-Agent Autonomous Reasoning in Hydrodynamics

Model ReleasesDGX agent

arXiv:2605.01102v1 Announce Type: new Abstract: Single-agent systems (SAS) have become the default pattern for LLM-driven scientific workflows, but routing planning, tool use, and synthesis through a

5 May 2026

Run an entire company with agents 🔥 It's always awesome to see companies continuing to innovate on AI-native UI/UX, particularly around mul…

AgentsDGX agent

Run an entire company with agents 🔥 It's always awesome to see companies continuing to innovate on AI-native UI/UX, particularly around multi-agent coordination, to solve deeply complex tasks beyond w

Semia: Auditing Agent Skills via Constraint-Guided Representation Synthesis

AgentsDGX agent

arXiv:2605.00314v1 Announce Type: cross Abstract: An agent skill is a configuration package that equips an LLM-driven agent with a concrete capability, such as reading email, executing shell commands,

SUDP: Secret-Use Delegation Protocol for Agentic Systems

AgentsDGX agent

arXiv:2604.24920v2 Announce Type: replace-cross Abstract: Agentic systems increasingly act with user secrets for APIs, messaging platforms, and cloud services. Today's bearer-secret interfaces impleme

4 May 2026

Build, Judge, Optimize: A Blueprint for Continuous Improvement of Multi-Agent Consumer Assistants

AgentsDGX agent

arXiv:2603.03565v2 Announce Type: replace-cross Abstract: Conversational shopping assistants (CSAs) represent a compelling application of agentic AI, but moving from prototype to production reveals tw

1 May 2026

Crab: A Semantics-Aware Checkpoint/Restore Runtime for Agent Sandboxes

AgentsDGX agent

arXiv:2604.28138v1 Announce Type: cross Abstract: Autonomous agents act through sandboxed containers and microVMs whose state spans filesystems, processes, and runtime artifacts. Checkpoint and restor

NEW paper from Microsoft Research. If you care about training computer-use agents, this is one to keep. (bookmark it) The team builds 1,000 …

AgentsDGX agent

NEW paper from Microsoft Research. If you care about training computer-use agents, this is one to keep. (bookmark it) The team builds 1,000 synthetic computers (each with realistic directory structure

shoutout to the man @kylejeong for cooking this up 🔥 every agent can now get a dedicated browser subagent in seconds to navigate sites, fil…

AgentsDGX agent

shoutout to the man @kylejeong for cooking this up 🔥 every agent can now get a dedicated browser subagent in seconds to navigate sites, fill out forms, click through things, you get it Deep Agents x B

30 Apr 2026

AGEL-Comp: A Neuro-Symbolic Framework for Compositional Generalization in Interactive Agents

AgentsDGX agent

arXiv:2604.26522v1 Announce Type: new Abstract: Large Language Model (LLM)-based agents exhibit systemic failures in compositional generalization, limiting their robustness in interactive environments

29 Apr 2026

DeepAgents Deploy is great, you get - a fully open agent harness (no hidden stuff, inspect everything) - managed infra - can use any cocktai…

AgentsDGX agent

DeepAgents Deploy is great, you get - a fully open agent harness (no hidden stuff, inspect everything) - managed infra - can use any cocktail of models including Open Models - and everything is traced

Exploring Reasoning Reward Model for Agents

Model ReleasesDGX agent

arXiv:2601.22154v2 Announce Type: replace-cross Abstract: Agentic Reinforcement Learning (Agentic RL) has achieved notable success in enabling agents to perform complex reasoning and tool use. However

This is so good @hwchase17 walks through the complete setup of a Deep Agents deployment step-by-step It’s a great example of getting started…

AgentsDGX agent

This is so good @hwchase17 walks through the complete setup of a Deep Agents deployment step-by-step It’s a great example of getting started fast with Deep Agents from the CLI! DeepAgents Deploy is th

28 Apr 2026

DRACULA: Hunting for the Actions Users Want Deep Research Agents to Execute

AgentsDGX agent

arXiv:2604.23815v1 Announce Type: new Abstract: Scientific Deep Research (DR) agents answer user queries by synthesizing research papers into multi-section reports. User feedback can improve their uti

Governing What You Cannot Observe: Adaptive Runtime Governance for Autonomous AI Agents

SafetyDGX agent

arXiv:2604.24686v1 Announce Type: new Abstract: Autonomous AI agents can remain fully authorized and still become unsafe as behavior drifts, adversaries adapt, and decision patterns shift without any

← Previous
1…3637383940…296
Next →