AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,153 results
Model Releases

ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning

DGX agent

arXiv:2605.20176v1 Announce Type: new Abstract: Large language models (LLMs) and agentic systems have shown promise for clinical decision support, but existing works largely assume that evidence has a

model-releasesarxiv-cs-cl
20 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Distribution-Free Uncertainty Quantification for Continuous AI Agent Evaluation

DGX agent

arXiv:2605.19779v1 Announce Type: new Abstract: We adapt split conformal prediction and adaptive conformal inference (ACI) to continuous AI agent evaluation, providing distribution-free coverage guara

model-releasesarxiv-cs-ai
20 May 2026
Local Ai

OEP: Poisoning Self-Evolving LLM Agents via Locally Correct but Non-Transferable Experiences

DGX agent

arXiv:2605.18930v1 Announce Type: cross Abstract: Memory-augmented large language model (LLM) agents use iterative reflection and self-evolution to solve complex tasks, but these mechanisms introduce

local-aiarxiv-cs-ai
20 May 2026
Safety

PEEK: Context Map as an Orientation Cache for Long-Context LLM Agents

DGX agent

arXiv:2605.19932v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly operate over long and recurring external contexts, like document corpora and code repositories. Across in

safetyarxiv-cs-ai
20 May 2026
Agents

Toward User Comprehension Supports for LLM Agent Skill Specifications

DGX agent

arXiv:2605.19362v1 Announce Type: cross Abstract: Users often interpret and select agent skills through their exttt{SKILL.md} specifications. To protect users, existing audits mainly focus on maliciou

agentsarxiv-cs-ai
20 May 2026
Model Releases

TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing

DGX agent

arXiv:2605.18859v1 Announce Type: cross Abstract: LLM routing matters most in long-horizon applications such as coding agents, deep research systems, and computer-use agents, where a single user reque

model-releasesarxiv-cs-ai
20 May 2026
Agents

What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents

DGX agent

arXiv:2605.19447v1 Announce Type: new Abstract: Reinforcement learning can train LLM agents from sparse task rewards, but long-horizon credit assignment remains challenging: a single success-or-failur

agentsarxiv-cs-ai
20 May 2026
Model Releases

Whispers of Wealth: Red-Teaming Google's Agent Payments Protocol via Prompt Injection

DGX agent

arXiv:2601.22569v2 Announce Type: replace-cross Abstract: Large language model (LLM) based agents are increasingly used to automate financial transactions, yet their reliance on contextual reasoning e

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Agentic AI Governance and Lifecycle Management in Healthcare

DGX agent

arXiv:2601.15630v2 Announce Type: replace Abstract: Healthcare organizations are beginning to embed agentic AI into routine workflows, including clinical documentation support and early-warning monito

model-releasesarxiv-cs-ai
19 May 2026
Safety

AMATA: Adaptive Multi-Agent Trajectory Alignment for Knowledge-Intensive Question Answering

DGX agent

arXiv:2605.17352v1 Announce Type: new Abstract: Despite substantial advances in large language models (LLMs), generating factually consistent responses for knowledge-intensive question answering remai

safetyarxiv-cs-cl
19 May 2026
Agents

Dynamic Generation of Multi-LLM Agents Communication Topologies with Graph Diffusion Models

DGX agent

arXiv:2510.07799v2 Announce Type: replace-cross Abstract: The efficiency of multi-agent systems driven by large language models (LLMs) largely hinges on their communication topology. However, designin

agentsarxiv-cs-ai
19 May 2026
Model Releases

Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps

DGX agent

arXiv:2605.17554v1 Announce Type: new Abstract: Frontier deep research agents (DRAs) plan a research task, synthesize across documents, and return a structured deliverable on demand. They are being de

model-releasesarxiv-cs-ai
19 May 2026
Agents

HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents

DGX agent

arXiv:2605.17873v1 Announce Type: cross Abstract: Training long-horizon LLM agents with reinforcement learning is challenging because sparse outcome rewards reveal whether a task succeeds, but not whi

agentsarxiv-cs-ai
19 May 2026
Safety

Learning Transferable Topology Priors for Multi-Agent LLM Collaboration Across Domains

DGX agent

arXiv:2605.17359v1 Announce Type: new Abstract: Large language model (LLM)-based multi-agent systems have shown strong potential for complex reasoning by coordinating specialized agents through struct

safetyarxiv-cs-cl
19 May 2026
Model Releases

SaaSBench: Exploring the Boundaries of Coding Agents in Long-Horizon Enterprise SaaS Engineering

DGX agent

arXiv:2605.17526v1 Announce Type: cross Abstract: As autonomous coding agents become capable of handling increasingly long-horizon tasks, they have gradually demonstrated the potential to complete end

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

SkillGenBench: Benchmarking Skill Generation Pipelines for LLM Agents

DGX agent

arXiv:2605.18693v1 Announce Type: new Abstract: As LLM agents are increasingly built around reusable skills, a central challenge is no longer only whether agents can use provided skills, but whether t

model-releasesarxiv-cs-ai
19 May 2026
Agents

Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception

DGX agent

arXiv:2601.09413v2 Announce Type: replace-cross Abstract: We introduce a voice-agentic framework that learns one critical omni-understanding skill: knowing when to trust itself versus when to consult

agentsarxiv-cs-ai
19 May 2026
Agents

Throughput-Optimal Scheduling Algorithms for LLM Inference and AI Agents

DGX agent

arXiv:2504.07347v3 Announce Type: replace-cross Abstract: As demand for Large Language Models (LLMs) and AI agents grows rapidly, optimizing systems for efficient LLM inference becomes critical. While

agentsarxiv-cs-lg
19 May 2026
Model Releases

Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design

DGX agent

arXiv:2605.15871v1 Announce Type: new Abstract: Toward recursive self-improvement, we investigate LLM agents autonomously designing foundation models beyond standard Transformers. We introduce a dual-

model-releasesarxiv-cs-ai
18 May 2026
Agents

Context Pruning for Coding Agents via Multi-Rubric Latent Reasoning

DGX agent

arXiv:2605.15315v1 Announce Type: new Abstract: LLM-powered coding agents spend the majority of their token budget reading repository files, yet much of the retrieved code is irrelevant to the task at

agentsarxiv-cs-ai
18 May 2026
Agents

DrugSAGE:Self-evolving Agent Experience for Efficient State-of-the-Art Drug Discovery

DGX agent

arXiv:2605.15461v1 Announce Type: cross Abstract: Building state-of-the-art (SOTA) predictive models for drug discovery requires expensive search over tools, architectures, and training strategies. Cu

agentsarxiv-cs-ai
18 May 2026
Agents

paper.json: A Coordination Convention for LLM-Agent-Actionable Papers

DGX agent

arXiv:2605.16194v1 Announce Type: cross Abstract: LLM agents routinely serve as first (and sometimes only) readers of academic papers, skimming for sub-claims, extracting reproducibility steps, and ge

agentsarxiv-cs-ai
18 May 2026
Agents

Runtime-Structured Task Decomposition for Agentic Coding Systems

DGX agent

arXiv:2605.15425v1 Announce Type: cross Abstract: Agentic coding systems increasingly use large language models (LLMs) for software engineering tasks such as debugging, root cause analysis, and code r

agentsarxiv-cs-ai
18 May 2026
Model Releases

Agentic Design of Compositional Descriptors via Autoresearch for Materials Science Applications

DGX agent

arXiv:2605.14671v1 Announce Type: cross Abstract: Autoresearch offers a flexible paradigm for automating scientific tasks, in which an AI agent proposes, implements, evaluates, and refines candidate s

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents

DGX agent

arXiv:2605.14133v1 Announce Type: new Abstract: Interactive agent benchmarks face a tension between scalable construction and realistic workflow evaluation. Hand-authored tasks are expensive to extend

model-releasesarxiv-cs-ai
15 May 2026
Agents

GEAR: Genetic AutoResearch for Agentic Code Evolution

DGX agent

arXiv:2605.13874v1 Announce Type: cross Abstract: Autonomous research agents can already run machine learning experiments without human supervision, but many rely on a narrow search strategy: they rep

agentsarxiv-cs-ai
15 May 2026
Model Releases

Herculean: An Agentic Benchmark for Financial Intelligence

DGX agent

arXiv:2605.14355v1 Announce Type: new Abstract: As AI agents improve, the central question is no longer whether they can solve isolated well-defined financial tasks, but whether they can reliably carr

model-releasesarxiv-cs-ai
15 May 2026
Agents

Lang2MLIP: End-to-End Language-to-Machine Learning Interatomic Potential Development with Autonomous Agentic Workflows

DGX agent

arXiv:2605.14527v1 Announce Type: new Abstract: Developing machine learning interatomic potentials (MLIPs) for complex materials systems remains challenging because it requires expertise in atomistic

agentsarxiv-cs-lg
15 May 2026
Agents

MemLineage: Lineage-Guided Enforcement for LLM Agent Memory

DGX agent

arXiv:2605.14421v1 Announce Type: cross Abstract: We introduce MemLineage, a defense for LLM agent memory that attaches both cryptographic provenance and LLM-mediated derivation lineage to every entry

agentsarxiv-cs-ai
15 May 2026
Model Releases

Polaris: A Godel Agent Framework for Small Language Models through Experience-Abstracted Policy Repair

DGX agent

arXiv:2603.23129v2 Announce Type: replace Abstract: Godel agent realize recursive self-improvement: an agent inspects its own policy and traces and then modifies that policy in a tested loop. We intro

model-releasesarxiv-cs-lg
15 May 2026
Safety

Progent: Securing AI Agents with Privilege Control

DGX agent

arXiv:2504.11703v3 Announce Type: replace-cross Abstract: AI agents interact with external environments through tool calls, exposing them to attacks like indirect prompt injection that can trigger una

safetyarxiv-cs-ai
15 May 2026
Agents

Remember Your Trace: Memory-Guided Long-Horizon Agentic Framework for Consistent and Hierarchical Repository-Level Code Documentation

DGX agent

arXiv:2605.14563v1 Announce Type: cross Abstract: Automated code documentation is essential for modern software development, providing the contextual grounding that both human developers and coding ag

agentsarxiv-cs-cl
15 May 2026
Model Releases

The Moltbook Observatory Archive: an incremental dataset of agent-only social network activity

DGX agent

arXiv:2605.13860v1 Announce Type: cross Abstract: Moltbook is a social media platform in which posts and comments are authored exclusively by autonomous AI agents. We present the Moltbook Observatory

model-releasesarxiv-cs-ai
15 May 2026
Agents

AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation

DGX agent

arXiv:2605.12925v1 Announce Type: cross Abstract: Evaluation of software engineering (SWE) agents is dominated by a binary signal: whether the final patch passes the tests. This outcome-only view trea

agentsarxiv-cs-ai
14 May 2026
Agents

An Agentic LLM-Based Framework for Population-Scale Mental Health Screening

DGX agent

arXiv:2605.13046v1 Announce Type: new Abstract: Mental health disorders affect millions worldwide, and healthcare systems are increasingly overwhelmed by the volume of clinical data generated from ele

agentsarxiv-cs-ai
14 May 2026
Agents

CADDesigner: Conceptual CAD Model Generation with a General-Purpose Agent

DGX agent

arXiv:2508.01031v5 Announce Type: replace Abstract: Computer-Aided Design (CAD) is widely used for conceptual design and parametric 3D modeling, but typically requires a high level of expertise from d

agentsarxiv-cs-ai
14 May 2026
Agents

Can LLM Agents Simulate Dynamic Networks? A Case Study on Email Networks with Phishing Synthesis

DGX agent

arXiv:2605.12507v1 Announce Type: cross Abstract: While Large Language Model (LLM) multi-agent systems (MAS) offer a transformative approach to simulating human behavior in complex systems, it remains

agentsarxiv-cs-ai
14 May 2026
Safety

Counterfactual Reasoning for Causal Responsibility Attribution in Probabilistic Multi-Agent Systems

DGX agent

arXiv:2605.13077v1 Announce Type: cross Abstract: Responsibility allocation -- determining the extent to which agents are accountable for outcomes -- is a fundamental challenge in the design and analy

safetyarxiv-cs-ai
14 May 2026
Agents

MAP: A Map-then-Act Paradigm for Long-Horizon Interactive Agent Reasoning

DGX agent

arXiv:2605.13037v1 Announce Type: new Abstract: Current interactive LLM agents rely on goal-conditioned stepwise planning, where environmental understanding is acquired reactively during execution rat

agentsarxiv-cs-ai
14 May 2026
Safety

No Attack Required: Semantic Fuzzing for Specification Violations in Agent Skills

DGX agent

arXiv:2605.13044v1 Announce Type: cross Abstract: LLM-powered agents can silently delete documents, leak credentials, or transfer funds on a routine user request, not because the agent was attacked, b

safetyarxiv-cs-ai
14 May 2026
Safety

Position: Assistive Agents Need Accessibility Alignment

DGX agent

arXiv:2605.13579v1 Announce Type: new Abstract: Assistive agents for Blind and Visually Impaired (BVI) users require accessibility alignment as a first-class design objective. Despite rapid progress i

safetyarxiv-cs-ai
14 May 2026
Safety

Robust and Safe Multi-Agent Reinforcement Learning with Communication for Autonomous Vehicles: From Simulation to Hardware

DGX agent

arXiv:2506.00982v3 Announce Type: replace Abstract: Deep multi-agent reinforcement learning (MARL) has been demonstrated effectively in simulations for multi-robot problems. For autonomous vehicles, t

safetyarxiv-cs-ro
14 May 2026
Model Releases

Toward Scalable Verifiable Reward: Proxy State-Based Evaluation for Multi-turn Tool-Calling LLM Agents

DGX agent

arXiv:2602.16246v3 Announce Type: replace Abstract: Interactive large language model (LLM) agents operating via multi-turn dialogue and multi-step tool calling are increasingly used in production. Ben

model-releasesarxiv-cs-ai
14 May 2026
Model Releases

When Does Hierarchy Help? Benchmarking Agent Coordination in Event-Driven Industrial Scheduling

DGX agent

arXiv:2605.13172v1 Announce Type: cross Abstract: Recent advances in agent and multi-agent systems have shown strong performance on tool use, reasoning, and collaborative tasks. However, existing benc

model-releasesarxiv-cs-ai
14 May 2026
Model Releases

ABRA: Agent Benchmark for Radiology Applications

DGX agent

arXiv:2605.11224v1 Announce Type: new Abstract: Existing medical-agent benchmarks deliver imaging as pre-selected samples, never as an environment the agent must navigate. We introduce ABRA, a radiolo

model-releasesarxiv-cs-cv
13 May 2026
Model Releases

GeomHerd: A Forward-looking Herding Quantification via Ricci Flow Geometry on Agent Interactive Simulations

DGX agent

arXiv:2605.11645v1 Announce Type: cross Abstract: Herding -- where agents align their behaviors and act collectively -- is a central driver of market fragility and systemic risk. Existing approaches t

model-releasesarxiv-cs-lg
13 May 2026
Agents

Multi-Agent System Identification with Nonlinear Sheaf Diffusion

DGX agent

arXiv:2605.11204v1 Announce Type: cross Abstract: Local interaction laws governing multi-agent systems can be difficult to recover from trajectory data, even when the dynamics are observed faithfully.

agentsarxiv-cs-lg
13 May 2026
Agents

On Problems of Implicit Context Compression for Software Engineering Agents

DGX agent

arXiv:2605.11051v1 Announce Type: cross Abstract: LLM-based Software Engineering agents face a critical bottleneck: context length limitations cause failures on complex, long-horizon tasks. One promis

agentsarxiv-cs-cl
13 May 2026
← Previous
1…5051525354…233
Next →