AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,958 results
1 Jun 2026

LongTraceRL: Learning Long-Context Reasoning from Search Agent Trajectories with Rubric Rewards

AgentsDGX agent

arXiv:2605.31584v1 Announce Type: cross Abstract: Long-context reasoning remains a central challenge for large language models, which often fail to locate and integrate key information in extensive di

MASPOB: Bandit-Based Prompt Optimization for Multi-Agent Systems with Graph Neural Networks

AgentsDGX agent

arXiv:2603.02630v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have achieved great success in many real-world applications, especially the one serving as the cognitive backbone

MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft

Model ReleasesDGX agent

arXiv:2605.30931v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and action generation. However, their ability to susta

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

NVIDIA Vera CPU Sets a New Standard for Agentic Workloads in AI Factories

HardwareDGX agent

NVIDIA Vera is a purpose-built CPU for agentic AI and reinforcement learning, delivering twice the efficiency and 50% faster performance than traditional rack-scale CPUs. The processor helps AI factor

the decade of agents! @theemozilla

HardwareDGX agent

the decade of agents! @theemozilla We have been working closely with @nvidia to ensure Hermes Agent works smoothly on their new @NVIDIARTXSpark superchip and integrates with the new OpenShell runtime,

29 May 2026

First head-to-head comparison of agentic AI applied to the analysis of simulated data of the Einstein Telescope

Model ReleasesDGX agent

arXiv:2605.28916v1 Announce Type: cross Abstract: We report a comparison of two state-of-the-art agentic AI systems, Claude Code (Anthropic) and Codex (OpenAI), tasked with autonomously executing a si

GTA: Generating Long-Horizon Tasks for Web Agents at Scale

Model ReleasesDGX agent

arXiv:2605.29218v1 Announce Type: new Abstract: Web agents, which couple language models with browsing and tool-use capabilities, show promise as open web assistants. Yet progress is increasingly limi

HEART-Bench: Do LLM Agents Exhibit Human-like Psychology?

Model ReleasesDGX agent

arXiv:2605.30058v1 Announce Type: new Abstract: While LLM agents have demonstrated remarkable task-oriented abilities such as planning, reasoning, and action, few works have treated them as complete h

KairosAgent: Agentic Time Series Forecasting with Fused Semantic Reasoning

AgentsDGX agent

arXiv:2605.30002v1 Announce Type: new Abstract: Cross-domain multimodal time series forecasting is a challenging task, requiring models to integrate precise numerical comprehension, cross-domain seman

MediHive: A Decentralized Agent Collective for Medical Reasoning

Local AiDGX agent

arXiv:2603.27150v2 Announce Type: replace Abstract: Large language models (LLMs) have revolutionized medical reasoning tasks, yet single-agent systems often falter on complex, interdisciplinary proble

Molecular Lead Optimization via Agentic Tool Planning

AgentsDGX agent

arXiv:2605.28862v1 Announce Type: new Abstract: Drug discovery is a lengthy and resource-intensive process composed of multiple stages. Among these stages, lead optimization plays a critical role in t

Redundant or Necessary? A Benchmark for Detecting Redundant Steps in Agent Trajectories

Model ReleasesDGX agent

arXiv:2605.29893v1 Announce Type: new Abstract: LLM-based agents have demonstrated strong capabilities in solving complex tasks through multi-step reasoning and tool use. However, existing evaluation

VFEAgent: A Multimodal Agent Framework for End-to-End Automated Finite Element Analysis

AgentsDGX agent

arXiv:2605.28978v1 Announce Type: new Abstract: Finite Element Analysis (FEA) serves as the cornerstone of modern engineering design. However, its workflow is inherently complex and relies heavily on

28 May 2026

A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks

Model ReleasesDGX agent

arXiv:2605.28556v1 Announce Type: new Abstract: As agent capabilities advance, existing benchmarks, such as au^2-Bench, are becoming increasingly saturated. Yet constructing new benchmark tasks remain

AgensFlow: A Coordination-Policy Substrate for Multi-Agent Systems

SafetyDGX agent

arXiv:2605.27466v1 Announce Type: cross Abstract: Multi-agent systems built on large language models (LLMs) require many coordination choices that are difficult to fix a priori: which skill protocol t

Agentic Separation Logic Specification Synthesis

Model ReleasesDGX agent

arXiv:2605.27531v1 Announce Type: cross Abstract: Specification synthesis, the task of automatically inferring formal specifications from program implementations and natural language, is important for

Automating Formal Verification with Agent-Guided Tree Search

Model ReleasesDGX agent

arXiv:2605.27485v1 Announce Type: cross Abstract: Formal verification offers a path to provably correct software, but writing verified code remains expensive enough that the technique is rarely used i

Build a test suite that grows with your agent with dataset management in Amazon Bedrock AgentCore

Model ReleasesDGX agent

Agent evaluation is most powerful when you combine fast-moving online signals with stable offline baselines. To understand whether your agent is truly improving over time, you need a fixed benchmark a

MIRAGE: Context-Aware Prompt Injection against Mobile GUI Agents via User-Generated Content

Model ReleasesDGX agent

arXiv:2605.28116v1 Announce Type: cross Abstract: Mobile graphical user interface (GUI) agents driven by vision-language models (VLMs) perceive the screen as rendered pixels and choose actions from wh

Multi-Agent LLM-based Metamorphic Testing for REST APIs

AgentsDGX agent

arXiv:2605.28321v1 Announce Type: cross Abstract: As REST APIs become an increasingly significant part of software systems, their validation is becoming more critical. Hence, testing and uncovering un

ReflexGrad: Within-Episode Failure Recovery in LLM Agents via Progress-Gated Dual-Process Routing

Model ReleasesDGX agent

arXiv:2511.14584v3 Announce Type: replace-cross Abstract: We present ReflexGrad, a dual-process architecture for within-episode failure recovery in LLM agents without demonstrations. When agents commi

ResearchMath-14K: Scaling Research-Level Mathematics via Agents

AgentsDGX agent

arXiv:2605.28003v1 Announce Type: new Abstract: The frontier of mathematics is defined by problems whose solutions are not yet known, yet it remains unclear whether language models can meaningfully en

Retrieval, Reward, and Training Protocols: What Matters in Training Search Agents?

SafetyDGX agent

arXiv:2605.27881v1 Announce Type: new Abstract: Search agents powered by large language models can autonomously decompose queries, retrieve information, and synthesize answers through multi-step reaso

SNARE: Adaptive Scenario Synthesis for Eliciting Overeager Behavior in Coding Agents

Model ReleasesDGX agent

arXiv:2605.28122v1 Announce Type: cross Abstract: A coding agent executes a benign task as a sequence of shell, file, and network actions, any of which can quietly exceed the authorized scope while th

Towards Faithful Agentic XAI: A Verification Method and an Open-World Benchmark for Better Model Faithfulness

Model ReleasesDGX agent

arXiv:2605.27879v1 Announce Type: new Abstract: Explainable AI (XAI) helps users interpret model behavior and identify potential faults. Agentic XAI systems use Large Language Models (LLMs) to make ex

TRACES: Proactive Safety Auditing for Multi-Turn LLM Agents via Trajectory-State Modeling

SafetyDGX agent

arXiv:2605.27690v1 Announce Type: new Abstract: LLM agents increasingly operate through multi-turn tool use and environment interaction, where safety risks often emerge from intermediate steps long be

27 May 2026

ChainCaps: Composition-Safe Tool-Using Agents via Monotonic Capability Attenuation

SafetyDGX agent

arXiv:2605.26542v1 Announce Type: cross Abstract: Tool-using agents increasingly operate in open-ended deployment environments, where they compose file systems, web APIs, code interpreters, and enterp

FinHarness: An Inline Lifecycle Safety Harness for Finance LLM Agents

SafetyDGX agent

arXiv:2605.27333v1 Announce Type: new Abstract: Finance LLM agents must simultaneously block prompt-induced unauthorized actions and approve legitimate multi-step business workflows. However, boundary

GENESIS: Harnessing AI Agents for Autonomous 6G RAN Synthesis, Research, and Testing

AgentsDGX agent

arXiv:2605.27360v1 Announce Type: cross Abstract: Cellular research and development (R&D) is throttled by six structural processes that each consume months of manual engineering work per iteration: (i

MobileExplorer: Accelerating On-Device Inference for Mobile GUI Agents via Online Exploration

Model ReleasesDGX agent

arXiv:2605.26546v1 Announce Type: new Abstract: Mobile graphical user interface (GUI) agents enable AI models to autonomously operate smartphones on behalf of users. However, most existing systems foc

SetupX: Can LLM Agents Learn from Past Failures in Functionality-Correct Code Repository Setup?

AgentsDGX agent

arXiv:2605.26186v1 Announce Type: cross Abstract: Functionality-correct repository setup aims to configure execution environments (e.g., dependencies, build scripts) to successfully execute a reposito

UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action

ApplicationsDGX agent

arXiv:2510.17790v3 Announce Type: replace-cross Abstract: Computer-use agents face a fundamental limitation. They rely exclusively on primitive GUI actions (click, type, scroll), creating brittle exec

26 May 2026

Anticipate and Learn: Unleashing Idle-Time Compute in Proactive Agents

Model ReleasesDGX agent

arXiv:2605.25971v1 Announce Type: new Abstract: While AI agents demonstrate remarkable capabilities in reasoning and tool use, they remain fundamentally reactive: they compute responses only after exp

Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows

Model ReleasesDGX agent

arXiv:2605.24219v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous agents that reason, use tools, and act over multiple steps. Yet most hallucination

Decoding ML Decision: An Agentic Reasoning Framework for Large-Scale Ranking System

AgentsDGX agent

arXiv:2602.18640v2 Announce Type: replace Abstract: Modern large-scale ranking systems operate within a sophisticated landscape of competing objectives, operational constraints, and evolving product r

Detectify debuts MCP server to let AI agents find and fix vulnerabilities in real time

AgentsDGX agent

Application security platform company Detectify AB today launched the Detectify MCP Server, a new integration layer that plugs the company’s security testing engines into artificial intelligence-drive

EvoSci: A Bio-Inspired Multi-Agent Framework for the Evolution of Scientific Discovery

AgentsDGX agent

arXiv:2605.24018v1 Announce Type: new Abstract: Large language models (LLMs), have shown strong potential in scientific discovery, yet existing methods still face substantial challenges in the design

GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning

Model ReleasesDGX agent

arXiv:2605.25200v1 Announce Type: new Abstract: Travel planning is a realistic task for evaluating the planning and tool-use abilities of LLM agents. However, existing benchmarks typically assume only

Hera: Learning Long-Horizon Coordination for Device-Cloud Collaborative LLM Agents

Local AiDGX agent

arXiv:2605.24598v1 Announce Type: new Abstract: Large language model (LLM) agents excel at solving complex long-horizon tasks through autonomous interaction with environments. However, their real-worl

IterInject: Indirect Prompt Injection Against LLM Agents via Feedback-Guided Iterative Optimization

Model ReleasesDGX agent

arXiv:2605.24659v1 Announce Type: new Abstract: LLM-based agents are increasingly deployed for complex tasks requiring planning, tool use, and interaction with external services. Their reliance on unt

LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries

Model ReleasesDGX agent

arXiv:2508.15760v2 Announce Type: replace-cross Abstract: Tool calling has emerged as a critical capability for AI agents. In contrast to conventional tool calling frameworks that rely on static, prov

MDIA: A Multi-Agent Diagnostic Intelligence Pipeline on HealthBench Professional

Model ReleasesDGX agent

arXiv:2605.24699v1 Announce Type: new Abstract: Most reported gains on agentic-LLM clinical benchmarks are often attributed to prompt engineering, yet our results suggest that larger improvements can

Multi-Agent Coordination Adaptation via Structure-Guided Orchestration

SafetyDGX agent

arXiv:2605.25746v1 Announce Type: cross Abstract: As large language model (LLM)-based multi-agent systems scale to handle increasingly complex tasks, balancing structural stability and dynamic adaptab

Multi-Agent Specification-based Metamorphic Testing of FMU-Based Simulations

AgentsDGX agent

arXiv:2605.25101v1 Announce Type: cross Abstract: In many industrial domains, the Functional Mock-up Interface (FMI) is used to exchange simulation models as Functional Mock-up Units (FMUs) across dif

Toward Reliable Design of LLM-Enabled Agentic Workflows: Optimizing Latency-Reliability-Cost Tradeoffs

SafetyDGX agent

arXiv:2605.23929v1 Announce Type: new Abstract: Modern AI systems increasingly rely on workflows composed of multiple interacting agents, some powered by large language models (LLMs) and others by con

When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs

SafetyDGX agent

arXiv:2605.24202v1 Announce Type: new Abstract: Multi-agent LLM workflows route inference through specialized roles to lift end-task accuracy, but jointly training those roles with reinforcement learn

25 May 2026

Agentic Proving for Program Verification

Model ReleasesDGX agent

arXiv:2605.23772v1 Announce Type: new Abstract: Agentic systems have recently emerged as state-of-the-art approaches for automated theorem proving in formal mathematics. To assess how far these capabi

Foundation Protocol: A Coordination Layer for Agentic Society

SafetyDGX agent

arXiv:2605.23218v1 Announce Type: new Abstract: Autonomous agents are moving from tools into a layer of social infrastructure: they browse, purchase, deploy software, manage systems, and increasingly

My favorite prompt: a) make a plan for <task> b) orchestrate and launch sub-agents to execute the plan c) validate the results from the sub-…

IndustryDGX agent

My favorite prompt: a) make a plan for <task> b) orchestrate and launch sub-agents to execute the plan c) validate the results from the sub-agents d) repeat b and c until you finish the plan Grok Buil

Parallel Context Compaction for Long-Horizon LLM Agent Serving

Model ReleasesDGX agent

arXiv:2605.23296v1 Announce Type: new Abstract: Long-horizon LLM agents accumulate growing conversation histories that eventually exceed the model's context window. Context compaction via LLM-based su

SkillOpt: Executive Strategy for Self-Evolving Agent Skills

Model ReleasesDGX agent

arXiv:2605.23904v1 Announce Type: new Abstract: Agent skills today are hand-crafted, generated one-shot, or evolved through loosely controlled self-revision, none of which behaves like a deep-learning

Whose Good, Whose Place? The Moral Geography of Agentic AI for Social Good

SafetyDGX agent

arXiv:2605.22995v1 Announce Type: cross Abstract: Agentic AI systems are increasingly proposed for social-good domains, often invoking the United Nations Sustainable Development Goals (SDGs) as a voca

23 May 2026

GraphFlow: A Graph-Based Workflow Management for Efficient LLM-Agent Serving

Model ReleasesDGX agent

arXiv:2605.22566v1 Announce Type: new Abstract: Large Language Model (LLM)-based agents demonstrate strong reasoning and execution capabilities on complex tasks when guided by structured instructions,

22 May 2026

Code Researcher: Deep Research Agent for Large Systems Code and Commit History

Model ReleasesDGX agent

arXiv:2506.11060v2 Announce Type: replace-cross Abstract: Large Language Model (LLM)-based coding agents have shown promising results on coding benchmarks, but their effectiveness on systems code rema

GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation

SafetyDGX agent

arXiv:2605.21605v1 Announce Type: new Abstract: Open-ended image generation is no longer a simple prompt-to-image problem. High-quality generation often requires an agent to combine a model's internal

MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks

Model ReleasesDGX agent

arXiv:2605.21917v1 Announce Type: new Abstract: Training Vision Language Models (VLMs) for video event reasoning requires high-quality structured annotations capturing not only what happened, but when

ProcBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents

Model ReleasesDGX agent

arXiv:2605.20251v2 Announce Type: cross Abstract: Existing benchmarks for LLM coding agents primarily evaluate final outcomes. While useful for measuring overall capability, these metrics provide limi

Tool-Augmented Agent for Closed-loop Optimization,Simulation,and Modeling Orchestration

AgentsDGX agent

arXiv:2605.20190v1 Announce Type: new Abstract: Iterative industrial design-simulation optimization is bottlenecked by the CAD-CAE semantic gap: translating simulation feedback into valid geometric ed

21 May 2026

1/ At @LangChain’s Interrupt conference last week, one question kept coming up: What happens when AI agents need to spend money? Enterprise …

ApplicationsDGX agent

1/ At @LangChain’s Interrupt conference last week, one question kept coming up: What happens when AI agents need to spend money? Enterprise agents have moved from prototype to production, but payments

Build AI agents for business intelligence with Amazon Bedrock AgentCore

Model ReleasesDGX agent

In this post, we show you how OPLOG developed three AI agents using the Strands Agents SDK, deployed them to Amazon Bedrock AgentCore, and integrated Amazon Bedrock with Anthropic’s Claude Sonnet and

← Previous
1…9596979899…300
Next →