AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,289 results
Model Releases

Redundant or Necessary? A Benchmark for Detecting Redundant Steps in Agent Trajectories

DGX agent

arXiv:2605.29893v1 Announce Type: new Abstract: LLM-based agents have demonstrated strong capabilities in solving complex tasks through multi-step reasoning and tool use. However, existing evaluation

model-releasesarxiv-cs-ai
29 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Agents

VFEAgent: A Multimodal Agent Framework for End-to-End Automated Finite Element Analysis

DGX agent

arXiv:2605.28978v1 Announce Type: new Abstract: Finite Element Analysis (FEA) serves as the cornerstone of modern engineering design. However, its workflow is inherently complex and relies heavily on

agentsarxiv-cs-ai
29 May 2026
Model Releases

A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks

DGX agent

arXiv:2605.28556v1 Announce Type: new Abstract: As agent capabilities advance, existing benchmarks, such as au^2-Bench, are becoming increasingly saturated. Yet constructing new benchmark tasks remain

model-releasesarxiv-cs-ai
28 May 2026
Safety

AgensFlow: A Coordination-Policy Substrate for Multi-Agent Systems

DGX agent

arXiv:2605.27466v1 Announce Type: cross Abstract: Multi-agent systems built on large language models (LLMs) require many coordination choices that are difficult to fix a priori: which skill protocol t

safetyarxiv-cs-ai
28 May 2026
Model Releases

Agentic Separation Logic Specification Synthesis

DGX agent

arXiv:2605.27531v1 Announce Type: cross Abstract: Specification synthesis, the task of automatically inferring formal specifications from program implementations and natural language, is important for

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Automating Formal Verification with Agent-Guided Tree Search

DGX agent

arXiv:2605.27485v1 Announce Type: cross Abstract: Formal verification offers a path to provably correct software, but writing verified code remains expensive enough that the technique is rarely used i

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

MIRAGE: Context-Aware Prompt Injection against Mobile GUI Agents via User-Generated Content

DGX agent

arXiv:2605.28116v1 Announce Type: cross Abstract: Mobile graphical user interface (GUI) agents driven by vision-language models (VLMs) perceive the screen as rendered pixels and choose actions from wh

model-releasesarxiv-cs-ai
28 May 2026
Agents

Multi-Agent LLM-based Metamorphic Testing for REST APIs

DGX agent

arXiv:2605.28321v1 Announce Type: cross Abstract: As REST APIs become an increasingly significant part of software systems, their validation is becoming more critical. Hence, testing and uncovering un

agentsarxiv-cs-ai
28 May 2026
Model Releases

ReflexGrad: Within-Episode Failure Recovery in LLM Agents via Progress-Gated Dual-Process Routing

DGX agent

arXiv:2511.14584v3 Announce Type: replace-cross Abstract: We present ReflexGrad, a dual-process architecture for within-episode failure recovery in LLM agents without demonstrations. When agents commi

model-releasesarxiv-cs-ai
28 May 2026
Agents

ResearchMath-14K: Scaling Research-Level Mathematics via Agents

DGX agent

arXiv:2605.28003v1 Announce Type: new Abstract: The frontier of mathematics is defined by problems whose solutions are not yet known, yet it remains unclear whether language models can meaningfully en

agentsarxiv-cs-cl
28 May 2026
Safety

Retrieval, Reward, and Training Protocols: What Matters in Training Search Agents?

DGX agent

arXiv:2605.27881v1 Announce Type: new Abstract: Search agents powered by large language models can autonomously decompose queries, retrieve information, and synthesize answers through multi-step reaso

safetyarxiv-cs-cl
28 May 2026
Model Releases

SNARE: Adaptive Scenario Synthesis for Eliciting Overeager Behavior in Coding Agents

DGX agent

arXiv:2605.28122v1 Announce Type: cross Abstract: A coding agent executes a benign task as a sequence of shell, file, and network actions, any of which can quietly exceed the authorized scope while th

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Towards Faithful Agentic XAI: A Verification Method and an Open-World Benchmark for Better Model Faithfulness

DGX agent

arXiv:2605.27879v1 Announce Type: new Abstract: Explainable AI (XAI) helps users interpret model behavior and identify potential faults. Agentic XAI systems use Large Language Models (LLMs) to make ex

model-releasesarxiv-cs-ai
28 May 2026
Safety

TRACES: Proactive Safety Auditing for Multi-Turn LLM Agents via Trajectory-State Modeling

DGX agent

arXiv:2605.27690v1 Announce Type: new Abstract: LLM agents increasingly operate through multi-turn tool use and environment interaction, where safety risks often emerge from intermediate steps long be

safetyarxiv-cs-cl
28 May 2026
Safety

ChainCaps: Composition-Safe Tool-Using Agents via Monotonic Capability Attenuation

DGX agent

arXiv:2605.26542v1 Announce Type: cross Abstract: Tool-using agents increasingly operate in open-ended deployment environments, where they compose file systems, web APIs, code interpreters, and enterp

safetyarxiv-cs-ai
27 May 2026
Safety

FinHarness: An Inline Lifecycle Safety Harness for Finance LLM Agents

DGX agent

arXiv:2605.27333v1 Announce Type: new Abstract: Finance LLM agents must simultaneously block prompt-induced unauthorized actions and approve legitimate multi-step business workflows. However, boundary

safetyarxiv-cs-cl
27 May 2026
Agents

GENESIS: Harnessing AI Agents for Autonomous 6G RAN Synthesis, Research, and Testing

DGX agent

arXiv:2605.27360v1 Announce Type: cross Abstract: Cellular research and development (R&D) is throttled by six structural processes that each consume months of manual engineering work per iteration: (i

agentsarxiv-cs-ai
27 May 2026
Model Releases

MobileExplorer: Accelerating On-Device Inference for Mobile GUI Agents via Online Exploration

DGX agent

arXiv:2605.26546v1 Announce Type: new Abstract: Mobile graphical user interface (GUI) agents enable AI models to autonomously operate smartphones on behalf of users. However, most existing systems foc

model-releasesarxiv-cs-ai
27 May 2026
Agents

SetupX: Can LLM Agents Learn from Past Failures in Functionality-Correct Code Repository Setup?

DGX agent

arXiv:2605.26186v1 Announce Type: cross Abstract: Functionality-correct repository setup aims to configure execution environments (e.g., dependencies, build scripts) to successfully execute a reposito

agentsarxiv-cs-ai
27 May 2026
Applications

UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action

DGX agent

arXiv:2510.17790v3 Announce Type: replace-cross Abstract: Computer-use agents face a fundamental limitation. They rely exclusively on primitive GUI actions (click, type, scroll), creating brittle exec

applicationsarxiv-cs-cl
27 May 2026
Model Releases

Anticipate and Learn: Unleashing Idle-Time Compute in Proactive Agents

DGX agent

arXiv:2605.25971v1 Announce Type: new Abstract: While AI agents demonstrate remarkable capabilities in reasoning and tool use, they remain fundamentally reactive: they compute responses only after exp

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows

DGX agent

arXiv:2605.24219v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous agents that reason, use tools, and act over multiple steps. Yet most hallucination

model-releasesarxiv-cs-ai
26 May 2026
Agents

Decoding ML Decision: An Agentic Reasoning Framework for Large-Scale Ranking System

DGX agent

arXiv:2602.18640v2 Announce Type: replace Abstract: Modern large-scale ranking systems operate within a sophisticated landscape of competing objectives, operational constraints, and evolving product r

agentsarxiv-cs-ai
26 May 2026
Agents

EvoSci: A Bio-Inspired Multi-Agent Framework for the Evolution of Scientific Discovery

DGX agent

arXiv:2605.24018v1 Announce Type: new Abstract: Large language models (LLMs), have shown strong potential in scientific discovery, yet existing methods still face substantial challenges in the design

agentsarxiv-cs-ai
26 May 2026
Model Releases

GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning

DGX agent

arXiv:2605.25200v1 Announce Type: new Abstract: Travel planning is a realistic task for evaluating the planning and tool-use abilities of LLM agents. However, existing benchmarks typically assume only

model-releasesarxiv-cs-cl
26 May 2026
Local Ai

Hera: Learning Long-Horizon Coordination for Device-Cloud Collaborative LLM Agents

DGX agent

arXiv:2605.24598v1 Announce Type: new Abstract: Large language model (LLM) agents excel at solving complex long-horizon tasks through autonomous interaction with environments. However, their real-worl

local-aiarxiv-cs-ai
26 May 2026
Model Releases

IterInject: Indirect Prompt Injection Against LLM Agents via Feedback-Guided Iterative Optimization

DGX agent

arXiv:2605.24659v1 Announce Type: new Abstract: LLM-based agents are increasingly deployed for complex tasks requiring planning, tool use, and interaction with external services. Their reliance on unt

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries

DGX agent

arXiv:2508.15760v2 Announce Type: replace-cross Abstract: Tool calling has emerged as a critical capability for AI agents. In contrast to conventional tool calling frameworks that rely on static, prov

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

MDIA: A Multi-Agent Diagnostic Intelligence Pipeline on HealthBench Professional

DGX agent

arXiv:2605.24699v1 Announce Type: new Abstract: Most reported gains on agentic-LLM clinical benchmarks are often attributed to prompt engineering, yet our results suggest that larger improvements can

model-releasesarxiv-cs-ai
26 May 2026
Safety

Multi-Agent Coordination Adaptation via Structure-Guided Orchestration

DGX agent

arXiv:2605.25746v1 Announce Type: cross Abstract: As large language model (LLM)-based multi-agent systems scale to handle increasingly complex tasks, balancing structural stability and dynamic adaptab

safetyarxiv-cs-ai
26 May 2026
Agents

Multi-Agent Specification-based Metamorphic Testing of FMU-Based Simulations

DGX agent

arXiv:2605.25101v1 Announce Type: cross Abstract: In many industrial domains, the Functional Mock-up Interface (FMI) is used to exchange simulation models as Functional Mock-up Units (FMUs) across dif

agentsarxiv-cs-ai
26 May 2026
Safety

Toward Reliable Design of LLM-Enabled Agentic Workflows: Optimizing Latency-Reliability-Cost Tradeoffs

DGX agent

arXiv:2605.23929v1 Announce Type: new Abstract: Modern AI systems increasingly rely on workflows composed of multiple interacting agents, some powered by large language models (LLMs) and others by con

safetyarxiv-cs-ai
26 May 2026
Safety

When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs

DGX agent

arXiv:2605.24202v1 Announce Type: new Abstract: Multi-agent LLM workflows route inference through specialized roles to lift end-task accuracy, but jointly training those roles with reinforcement learn

safetyarxiv-cs-ai
26 May 2026
Model Releases

Agentic Proving for Program Verification

DGX agent

arXiv:2605.23772v1 Announce Type: new Abstract: Agentic systems have recently emerged as state-of-the-art approaches for automated theorem proving in formal mathematics. To assess how far these capabi

model-releasesarxiv-cs-ai
25 May 2026
Safety

Foundation Protocol: A Coordination Layer for Agentic Society

DGX agent

arXiv:2605.23218v1 Announce Type: new Abstract: Autonomous agents are moving from tools into a layer of social infrastructure: they browse, purchase, deploy software, manage systems, and increasingly

safetyarxiv-cs-ai
25 May 2026
Model Releases

Parallel Context Compaction for Long-Horizon LLM Agent Serving

DGX agent

arXiv:2605.23296v1 Announce Type: new Abstract: Long-horizon LLM agents accumulate growing conversation histories that eventually exceed the model's context window. Context compaction via LLM-based su

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

SkillOpt: Executive Strategy for Self-Evolving Agent Skills

DGX agent

arXiv:2605.23904v1 Announce Type: new Abstract: Agent skills today are hand-crafted, generated one-shot, or evolved through loosely controlled self-revision, none of which behaves like a deep-learning

model-releasesarxiv-cs-ai
25 May 2026
Safety

Whose Good, Whose Place? The Moral Geography of Agentic AI for Social Good

DGX agent

arXiv:2605.22995v1 Announce Type: cross Abstract: Agentic AI systems are increasingly proposed for social-good domains, often invoking the United Nations Sustainable Development Goals (SDGs) as a voca

safetyarxiv-cs-ai
25 May 2026
Model Releases

GraphFlow: A Graph-Based Workflow Management for Efficient LLM-Agent Serving

DGX agent

arXiv:2605.22566v1 Announce Type: new Abstract: Large Language Model (LLM)-based agents demonstrate strong reasoning and execution capabilities on complex tasks when guided by structured instructions,

model-releasesarxiv-cs-lg
23 May 2026
Model Releases

Code Researcher: Deep Research Agent for Large Systems Code and Commit History

DGX agent

arXiv:2506.11060v2 Announce Type: replace-cross Abstract: Large Language Model (LLM)-based coding agents have shown promising results on coding benchmarks, but their effectiveness on systems code rema

model-releasesarxiv-cs-ai
22 May 2026
Safety

GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation

DGX agent

arXiv:2605.21605v1 Announce Type: new Abstract: Open-ended image generation is no longer a simple prompt-to-image problem. High-quality generation often requires an agent to combine a model's internal

safetyarxiv-cs-cv
22 May 2026
Model Releases

MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks

DGX agent

arXiv:2605.21917v1 Announce Type: new Abstract: Training Vision Language Models (VLMs) for video event reasoning requires high-quality structured annotations capturing not only what happened, but when

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

ProcBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents

DGX agent

arXiv:2605.20251v2 Announce Type: cross Abstract: Existing benchmarks for LLM coding agents primarily evaluate final outcomes. While useful for measuring overall capability, these metrics provide limi

model-releasesarxiv-cs-ai
22 May 2026
Agents

Tool-Augmented Agent for Closed-loop Optimization,Simulation,and Modeling Orchestration

DGX agent

arXiv:2605.20190v1 Announce Type: new Abstract: Iterative industrial design-simulation optimization is bottlenecked by the CAD-CAE semantic gap: translating simulation feedback into valid geometric ed

agentsarxiv-cs-ai
22 May 2026
Model Releases

FT-Dojo: Towards Autonomous LLM Fine-Tuning with Language Agents

DGX agent

arXiv:2603.01712v2 Announce Type: replace-cross Abstract: Fine-tuning large language models for vertical domains remains labor-intensive, requiring practitioners to curate data, configure training, an

model-releasesarxiv-cs-lg
21 May 2026
Agents

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning

DGX agent

arXiv:2605.20342v1 Announce Type: new Abstract: Training large multimodal models (LMMs) via reinforcement learning (RL) to natively invoke video-processing tools (e.g., cropping) has become a promisin

agentsarxiv-cs-cv
21 May 2026
Model Releases

EngiAI: A Multi-Agent Framework and Benchmark Suite for LLM-Driven Engineering Design

DGX agent

arXiv:2605.19743v1 Announce Type: new Abstract: Large Language Model (LLM) agents are increasingly applied to engineering design tasks, yet existing evaluation frameworks do not adequately address mul

model-releasesarxiv-cs-ai
20 May 2026
Agents

From Intent to AI Pipelines: A Controlled Agentic Framework for Non-AI Expert Scientists

DGX agent

arXiv:2605.18764v1 Announce Type: cross Abstract: Artificial Intelligence (AI) pipelines have become integral to modern research, supporting fields such as Medical Sciences, Agriculture, and Social Sc

agentsarxiv-cs-ai
20 May 2026
← Previous
1…6970717273…236
Next →