AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,153 results
Model Releases

Demonstration-Free Robotic Control via LLM Agents

DGX agent

arXiv:2601.20334v2 Announce Type: replace-cross Abstract: Robotic manipulation has increasingly adopted vision-language-action (VLA) models, which achieve strong performance but typically require task

model-releasesarxiv-cs-ai
30 Jun 2026
Agents

Experience Graphs: The Data Foundation for Self-Improving Agents

AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2606.29823v1 Announce Type: cross Abstract: The database community has repeatedly advanced the state of the art by recognizing that new workloads demand new system architectures. We argue that l

agentsarxiv-cs-ai
30 Jun 2026
Model Releases

Multi-Agent Routing as Set-Valued Prediction: A WildChat Benchmark and Cost-Aware Evaluation

DGX agent

arXiv:2606.28925v1 Announce Type: cross Abstract: Tool and agent routing from natural-language prompts is naturally a set-valued prediction problem: a single query may require multiple agents, while o

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

SEATauBench: Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages

DGX agent

arXiv:2606.28715v1 Announce Type: cross Abstract: While AI development and evaluation for Southeast Asia (SEA) has grown rapidly, agent capabilities in regional languages are still poorly understood d

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

TraceLab: Characterizing Coding Agent Workloads for LLM Serving

DGX agent

arXiv:2606.30560v1 Announce Type: cross Abstract: Coding agents are rapidly becoming a major application of agentic LLMs, but serving them efficiently remains challenging. Progress on this challenge r

model-releasesarxiv-cs-ai
30 Jun 2026
Agents

What Capable Agents Must Know: Selection Theorems for Robust Decision-Making under Uncertainty

DGX agent

arXiv:2603.02491v3 Announce Type: replace-cross Abstract: As artificial agents become increasingly capable, what internal structure is necessary for an agent to act competently under uncertainty? Clas

agentsarxiv-cs-ai
30 Jun 2026
Model Releases

Contagion Networks: Evaluator Preference Propagation in Multi-Agent LLM Systems

DGX agent

arXiv:2606.20493v2 Announce Type: replace-cross Abstract: When large language models serve as evaluators in multi-agent systems, their strategy preferences -- whether induced by explicit prompts or by

model-releasesarxiv-cs-ai
29 Jun 2026
Agents

Hierarchical Control in Multi-Agent Games: LLM-based Planning and RL Execution

DGX agent

arXiv:2606.20014v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has achieved strong performance in sequential decision-making, yet scaling to complex multi-agent environments rem

agentsarxiv-cs-ai
29 Jun 2026
Model Releases

LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior

DGX agent

arXiv:2606.28182v1 Announce Type: cross Abstract: Embodied agents operating in decentralized and partially observable environments have attracted growing attention in recent years. However, existing l

model-releasesarxiv-cs-ai
29 Jun 2026
Model Releases

The Red Queen Godel Machine: Co-Evolving Agents and Their Evaluators

DGX agent

arXiv:2606.26294v1 Announce Type: cross Abstract: Self-improving agents are state-of-the-art (SOTA) on agentic coding benchmarks and have recently been extended to general domains. However, their sear

model-releasesarxiv-cs-ai
26 Jun 2026
Agents

Diagnosing and Mitigating Compounding Failures in Agentic Persuasion via Taxonomic Strategy Retrieval

DGX agent

arXiv:2606.24976v1 Announce Type: cross Abstract: Foundation-model agents in multi-step, open-ended environments frequently suffer from compounding errors, where early mistakes contaminate long-horizo

agentsarxiv-cs-cl
25 Jun 2026
Agents

Debate2Create: Robot Co-design via Multi-Agent LLM Debate

DGX agent

arXiv:2510.25850v3 Announce Type: replace-cross Abstract: We introduce Debate2Create (D2C), a multi-agent LLM framework that formulates robot co-design as structured, iterative debate grounded in phys

agentsarxiv-cs-lg
24 Jun 2026
Safety

Memory Contagion: Cross-Temporal Propagation of Evaluator Bias via Agent Memory

DGX agent

arXiv:2606.23195v1 Announce Type: new Abstract: Large Language Model (LLM) agents increasingly rely on memory systems to maintain long-term coherence. Recent work shows that agent memories degrade dur

safetyarxiv-cs-lg
23 Jun 2026
Model Releases

Steer, Don't Solve: Training Small Critic Models for Large Code Agents

DGX agent

arXiv:2606.21811v1 Announce Type: cross Abstract: End-to-end code agent training is resource-intensive and plateaus on the strategy-level reasoning needed to resolve code issues, since jointly optimiz

model-releasesarxiv-cs-lg
23 Jun 2026
Model Releases

AI Coding Agents Can Reproduce Social Science Findings

DGX agent

arXiv:2606.11447v1 Announce Type: new Abstract: Recent anecdotal evidence suggests that AI coding agents can reproduce published findings when provided with original data and code; yet systematic eval

model-releasesarxiv-cs-cl
11 Jun 2026
Agents

Knowing When to Ask: Self-Gated Clarification for Hierarchical Language Agents

DGX agent

arXiv:2606.11349v1 Announce Type: new Abstract: In hierarchical reasoning, failures often originate at intermediate decision points where the agent commits to a wrong branch without recognizing that i

agentsarxiv-cs-ai
11 Jun 2026
Agents

Organize then Retrieve: Hierarchical Memory Navigation for Efficient Agents

DGX agent

arXiv:2606.11680v1 Announce Type: new Abstract: Large language model (LLM) agents struggle with long-horizon tasks due to their inherent statelessness, requiring all task-relevant information to be en

agentsarxiv-cs-ai
11 Jun 2026
Agents

The Confident Liar: Diagnosing Multi-Agent Debate with Log-Probabilities and LLM-as-Judge

DGX agent

arXiv:2606.10296v1 Announce Type: cross Abstract: Multi-agent debate systems are typically evaluated only on whether the final answer is correct, overlooking the quality of the intermediate reasoning

agentsarxiv-cs-ai
10 Jun 2026
Agents

VISTA: A Versatile Interactive User Simulation Toolkit for Agent Evaluation

DGX agent

arXiv:2606.11079v1 Announce Type: new Abstract: Evaluation remains a critical bottleneck for interactive agent development. Existing evaluation methods often rely on static benchmarks, which fail to c

agentsarxiv-cs-cl
10 Jun 2026
Agents

MetaEvo: A Meta-Optimization Framework for Experience-Driven Agent Evolution

DGX agent

arXiv:2606.07603v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit strong reasoning capabilities, yet most LLM-based agents are statically deployed and unable to improve through ta

agentsarxiv-cs-ai
9 Jun 2026
Agents

PACE: Anytime-Valid Acceptance Tests for Self-Evolving Agents

DGX agent

arXiv:2606.08106v1 Announce Type: new Abstract: Self-evolving agents improve by repeatedly proposing changes to their own prompts, skills, or workflows and keeping those that score higher on a small h

agentsarxiv-cs-ai
9 Jun 2026
Model Releases

SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?

DGX agent

arXiv:2606.07682v1 Announce Type: cross Abstract: AI agents are increasingly expected to complete long-horizon workflows that require sustained progress over hours, millions of tokens, and complex env

model-releasesarxiv-cs-ai
9 Jun 2026
Agents

The Token Not Taken: Sampling, State, and the Variability of AI Agent Outputs

DGX agent

arXiv:2606.08998v1 Announce Type: new Abstract: Agentic AI systems can behave differently across runs: the same request may produce a different plan, a different tool call, a different code edit, or a

agentsarxiv-cs-ai
9 Jun 2026
Agents

AdMem: Advanced Memory for Task-solving Agents

DGX agent

arXiv:2606.06787v1 Announce Type: new Abstract: Large Language Models (LLMs) show promise as tool-using agents but remain limited in long-horizon tasks that require remembering, organizing, and reusin

agentsarxiv-cs-ai
8 Jun 2026
Agents

AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning

DGX agent

arXiv:2512.13278v2 Announce Type: replace Abstract: Agentic reinforcement learning has advanced large language models (LLMs) to reason through long chain-of-thought trajectories while interleaving ext

agentsarxiv-cs-cl
8 Jun 2026
Model Releases

Declarative Skills for AI Agents in Knowledge-Grounded Tool-Use Workflows

DGX agent

arXiv:2606.06923v1 Announce Type: new Abstract: We study orchestration mechanisms for tool-using AI agents in realistic customer-service workflows over an unstructured knowledge base. We argue that de

model-releasesarxiv-cs-ai
8 Jun 2026
Agents

The Sim-to-Real Gap of Foundation Model Agents: A Unified MDP Perspective

DGX agent

arXiv:2606.07017v1 Announce Type: new Abstract: Foundation model agents are increasingly deployed for real-world decision-making, but suffer from the sim-to-real gap. While robotics and classical cont

agentsarxiv-cs-ai
8 Jun 2026
Safety

AdaMEM: Test-Time Adaptive Memory for Language Agents

DGX agent

arXiv:2606.05684v1 Announce Type: new Abstract: A central challenge for language agents is utilizing past experience to adapt to dynamic test-time conditions. While recent work demonstrates the promis

safetyarxiv-cs-ai
6 Jun 2026
Agents

Beyond Semantic Organization: Memory as Execution State Management for Long-Horizon Agents

DGX agent

arXiv:2606.06090v1 Announce Type: new Abstract: LLM-based agents increasingly tackle long-horizon tasks with interdependent decisions, where each action reshapes future constraints and intermediate er

agentsarxiv-cs-ai
6 Jun 2026
Agents

Insurance of Agentic AI

DGX agent

arXiv:2606.05449v1 Announce Type: new Abstract: Agentic artificial intelligence (AI) systems are transforming the risk landscape by extending beyond information generation to autonomous planning, tool

agentsarxiv-cs-ai
6 Jun 2026
Agents

Knowledge Activation: AI Skills as the Institutional Knowledge Primitive for Agentic Software Development

DGX agent

arXiv:2603.14805v2 Announce Type: replace Abstract: Enterprise software organizations accumulate critical institutional knowledge - architectural decisions, deployment procedures, compliance policies,

agentsarxiv-cs-ai
6 Jun 2026
Agents

Personal AI Agent for Camera Roll VQA

DGX agent

arXiv:2606.05275v1 Announce Type: new Abstract: We study the personal camera roll visual question answering setting. In this setting, a conversational AI assistant can access a user's personal camera

agentsarxiv-cs-cv
5 Jun 2026
Agents

Retrospective Harness Optimization: Improving LLM Agents via Self-Preference over Trajectory Rollouts

DGX agent

arXiv:2606.05922v1 Announce Type: cross Abstract: AI agents rely on a harness of skills, tools, and workflows to solve complex problems. Continually improving this harness is essential for adapting to

agentsarxiv-cs-cl
5 Jun 2026
Model Releases

AIP: A Graph Representation for Learning and Governing Agent Skills

DGX agent

arXiv:2606.04781v1 Announce Type: new Abstract: Agent Skills today consist largely of free-form prose requiring the agent to read, interpret, and re-derive how to act in every session. This imposes tw

model-releasesarxiv-cs-ai
4 Jun 2026
Agents

Beyond Prompt-Based Planning: MCP-Native Graph Planning-based Biomedical Agent System

DGX agent

arXiv:2606.04494v1 Announce Type: new Abstract: Biomedical agents promise to automate complex biological workflows, yet current systems face two fundamental bottlenecks: bioinformatics tools are highl

agentsarxiv-cs-ai
4 Jun 2026
Model Releases

Exploring the Topology and Memory of Consensus: How LLM Agents Agree, Fragment, or Settle When Forming Conventions

DGX agent

arXiv:2606.04197v1 Announce Type: cross Abstract: How much should an LLM agent remember, and how should multi-agent systems be connected when trying to reach consensus? We show these two design choice

model-releasesarxiv-cs-cl
4 Jun 2026
Agents

MIRAGE: Mobile Agents with Implicit Reasoning and Generative World Models

DGX agent

arXiv:2606.04627v1 Announce Type: new Abstract: Mobile agents are increasingly expected to operate everyday applications from screenshots and language goals, where reliable control requires reasoning

agentsarxiv-cs-ai
4 Jun 2026
Agents

Strabo: Declarative Specification and Implementation of Agentic Interaction Protocols

DGX agent

arXiv:2606.05043v1 Announce Type: new Abstract: The last few years have witnessed major advances in the modeling and implementation of multiagent systems based on declarative interaction protocols. Ou

agentsarxiv-cs-ai
4 Jun 2026
Agents

Topology Matters: Measuring Memory Leakage in Multi-Agent LLMs

DGX agent

arXiv:2512.04668v4 Announce Type: replace-cross Abstract: Graph topology is a fundamental determinant of memory leakage in multi-agent LLM systems, yet its effects remain poorly quantified. We introdu

agentsarxiv-cs-ai
4 Jun 2026
Agents

Automata-Conditioned Cooperative Multi-Agent Reinforcement Learning

DGX agent

arXiv:2511.02304v2 Announce Type: replace-cross Abstract: We study learning multi-task, multi-agent policies for cooperative, temporal objectives, under centralized training, decentralized execution.

agentsarxiv-cs-ai
3 Jun 2026
Model Releases

RealClawBench: Live OpenClaw Benchmarks from Real Developer-Agent Sessions

DGX agent

arXiv:2606.03889v1 Announce Type: new Abstract: Agent benchmarks should reflect what users actually ask deployed agents to do, yet existing benchmarks often miss key realism properties of real develop

model-releasesarxiv-cs-cl
3 Jun 2026
Agents

ToolGate: Token-Efficient Pre-Call Control for Tool-Augmented Vision-Language Agents

DGX agent

arXiv:2606.03054v1 Announce Type: new Abstract: Tool-augmented vision-language agents can acquire external perceptual evidence through OCR, detection, segmentation, and other tools, but executing ever

agentsarxiv-cs-ai
3 Jun 2026
Model Releases

WRIT: Write-Read Intensive Trajectory Synthesis for Multi-Turn User-Facing Agents

DGX agent

arXiv:2606.02908v1 Announce Type: cross Abstract: Multi-turn user-facing agents must infer user intent from incomplete requests, collect missing information through dialogue and tools, and execute val

model-releasesarxiv-cs-ai
3 Jun 2026
Agents

Agentic Authoring of Interactive Multiview Visualizations in Genomics

DGX agent

arXiv:2606.00370v1 Announce Type: cross Abstract: Diverse genomics data, scientific questions, and analysis tasks typically demand highly specialized visualizations. Therefore, users often must custom

agentsarxiv-cs-ai
2 Jun 2026
Agents

AgentxGCore: Agentic AI for Next-Generation Mobile Core Network

DGX agent

arXiv:2606.00417v1 Announce Type: cross Abstract: To meet the stringent requirements of emerging applications and the increasingly complex network management and operation, the Next Generation Mobile

agentsarxiv-cs-ai
2 Jun 2026
Agents

Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?

DGX agent

arXiv:2603.03202v3 Announce Type: replace Abstract: As large language models (LLMs) advance their mathematical capabilities toward the IMO and research level, the scarcity of challenging, high-quality

agentsarxiv-cs-cl
2 Jun 2026
Agents

Critic-R: Improving Agentic Search using Instruction-tuned Retrievers with Natural Language Introspective Feedback

DGX agent

arXiv:2606.00590v1 Announce Type: cross Abstract: Agentic search systems iteratively interact with retrieval models to answer complex queries. Despite substantial progress, optimizing retrievers for a

agentsarxiv-cs-ai
2 Jun 2026
Agents

Defenses & Enablers For Skill Injection Attacks on Terminal Based Agents

DGX agent

arXiv:2606.01567v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly rely on reusable skills i.e. documents describing task-specific procedures. However, this introduces a

agentsarxiv-cs-ai
2 Jun 2026
← Previous
1…3233343536…233
Next →