AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,959 results
Agents

LongTraceRL: Learning Long-Context Reasoning from Search Agent Trajectories with Rubric Rewards

DGX agent

arXiv:2605.31584v1 Announce Type: cross Abstract: Long-context reasoning remains a central challenge for large language models, which often fail to locate and integrate key information in extensive di

agentsarxiv-cs-ai
1 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Agents

MASPOB: Bandit-Based Prompt Optimization for Multi-Agent Systems with Graph Neural Networks

DGX agent

arXiv:2603.02630v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have achieved great success in many real-world applications, especially the one serving as the cognitive backbone

agentsarxiv-cs-ai
1 Jun 2026
Model Releases

MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft

DGX agent

arXiv:2605.30931v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and action generation. However, their ability to susta

model-releasesarxiv-cs-cl
1 Jun 2026
Hardware

NVIDIA Vera CPU Sets a New Standard for Agentic Workloads in AI Factories

DGX agent

NVIDIA Vera is a purpose-built CPU for agentic AI and reinforcement learning, delivering twice the efficiency and 50% faster performance than traditional rack-scale CPUs. The processor helps AI factor

hardwarenvidia-developer
1 Jun 2026
Hardware

the decade of agents! @theemozilla

DGX agent

the decade of agents! @theemozilla We have been working closely with @nvidia to ensure Hermes Agent works smoothly on their new @NVIDIARTXSpark superchip and integrates with the new OpenShell runtime,

hardwarenous-research--x
1 Jun 2026
Model Releases

First head-to-head comparison of agentic AI applied to the analysis of simulated data of the Einstein Telescope

DGX agent

arXiv:2605.28916v1 Announce Type: cross Abstract: We report a comparison of two state-of-the-art agentic AI systems, Claude Code (Anthropic) and Codex (OpenAI), tasked with autonomously executing a si

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

GTA: Generating Long-Horizon Tasks for Web Agents at Scale

DGX agent

arXiv:2605.29218v1 Announce Type: new Abstract: Web agents, which couple language models with browsing and tool-use capabilities, show promise as open web assistants. Yet progress is increasingly limi

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

HEART-Bench: Do LLM Agents Exhibit Human-like Psychology?

DGX agent

arXiv:2605.30058v1 Announce Type: new Abstract: While LLM agents have demonstrated remarkable task-oriented abilities such as planning, reasoning, and action, few works have treated them as complete h

model-releasesarxiv-cs-cl
29 May 2026
Agents

KairosAgent: Agentic Time Series Forecasting with Fused Semantic Reasoning

DGX agent

arXiv:2605.30002v1 Announce Type: new Abstract: Cross-domain multimodal time series forecasting is a challenging task, requiring models to integrate precise numerical comprehension, cross-domain seman

agentsarxiv-cs-ai
29 May 2026
Local Ai

MediHive: A Decentralized Agent Collective for Medical Reasoning

DGX agent

arXiv:2603.27150v2 Announce Type: replace Abstract: Large language models (LLMs) have revolutionized medical reasoning tasks, yet single-agent systems often falter on complex, interdisciplinary proble

local-aiarxiv-cs-ai
29 May 2026
Agents

Molecular Lead Optimization via Agentic Tool Planning

DGX agent

arXiv:2605.28862v1 Announce Type: new Abstract: Drug discovery is a lengthy and resource-intensive process composed of multiple stages. Among these stages, lead optimization plays a critical role in t

agentsarxiv-cs-lg
29 May 2026
Model Releases

Redundant or Necessary? A Benchmark for Detecting Redundant Steps in Agent Trajectories

DGX agent

arXiv:2605.29893v1 Announce Type: new Abstract: LLM-based agents have demonstrated strong capabilities in solving complex tasks through multi-step reasoning and tool use. However, existing evaluation

model-releasesarxiv-cs-ai
29 May 2026
Agents

VFEAgent: A Multimodal Agent Framework for End-to-End Automated Finite Element Analysis

DGX agent

arXiv:2605.28978v1 Announce Type: new Abstract: Finite Element Analysis (FEA) serves as the cornerstone of modern engineering design. However, its workflow is inherently complex and relies heavily on

agentsarxiv-cs-ai
29 May 2026
Model Releases

A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks

DGX agent

arXiv:2605.28556v1 Announce Type: new Abstract: As agent capabilities advance, existing benchmarks, such as au^2-Bench, are becoming increasingly saturated. Yet constructing new benchmark tasks remain

model-releasesarxiv-cs-ai
28 May 2026
Safety

AgensFlow: A Coordination-Policy Substrate for Multi-Agent Systems

DGX agent

arXiv:2605.27466v1 Announce Type: cross Abstract: Multi-agent systems built on large language models (LLMs) require many coordination choices that are difficult to fix a priori: which skill protocol t

safetyarxiv-cs-ai
28 May 2026
Model Releases

Agentic Separation Logic Specification Synthesis

DGX agent

arXiv:2605.27531v1 Announce Type: cross Abstract: Specification synthesis, the task of automatically inferring formal specifications from program implementations and natural language, is important for

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Automating Formal Verification with Agent-Guided Tree Search

DGX agent

arXiv:2605.27485v1 Announce Type: cross Abstract: Formal verification offers a path to provably correct software, but writing verified code remains expensive enough that the technique is rarely used i

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

Build a test suite that grows with your agent with dataset management in Amazon Bedrock AgentCore

DGX agent

Agent evaluation is most powerful when you combine fast-moving online signals with stable offline baselines. To understand whether your agent is truly improving over time, you need a fixed benchmark a

model-releasesaws-ml-blog
28 May 2026
Model Releases

MIRAGE: Context-Aware Prompt Injection against Mobile GUI Agents via User-Generated Content

DGX agent

arXiv:2605.28116v1 Announce Type: cross Abstract: Mobile graphical user interface (GUI) agents driven by vision-language models (VLMs) perceive the screen as rendered pixels and choose actions from wh

model-releasesarxiv-cs-ai
28 May 2026
Agents

Multi-Agent LLM-based Metamorphic Testing for REST APIs

DGX agent

arXiv:2605.28321v1 Announce Type: cross Abstract: As REST APIs become an increasingly significant part of software systems, their validation is becoming more critical. Hence, testing and uncovering un

agentsarxiv-cs-ai
28 May 2026
Model Releases

ReflexGrad: Within-Episode Failure Recovery in LLM Agents via Progress-Gated Dual-Process Routing

DGX agent

arXiv:2511.14584v3 Announce Type: replace-cross Abstract: We present ReflexGrad, a dual-process architecture for within-episode failure recovery in LLM agents without demonstrations. When agents commi

model-releasesarxiv-cs-ai
28 May 2026
Agents

ResearchMath-14K: Scaling Research-Level Mathematics via Agents

DGX agent

arXiv:2605.28003v1 Announce Type: new Abstract: The frontier of mathematics is defined by problems whose solutions are not yet known, yet it remains unclear whether language models can meaningfully en

agentsarxiv-cs-cl
28 May 2026
Safety

Retrieval, Reward, and Training Protocols: What Matters in Training Search Agents?

DGX agent

arXiv:2605.27881v1 Announce Type: new Abstract: Search agents powered by large language models can autonomously decompose queries, retrieve information, and synthesize answers through multi-step reaso

safetyarxiv-cs-cl
28 May 2026
Model Releases

SNARE: Adaptive Scenario Synthesis for Eliciting Overeager Behavior in Coding Agents

DGX agent

arXiv:2605.28122v1 Announce Type: cross Abstract: A coding agent executes a benign task as a sequence of shell, file, and network actions, any of which can quietly exceed the authorized scope while th

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Towards Faithful Agentic XAI: A Verification Method and an Open-World Benchmark for Better Model Faithfulness

DGX agent

arXiv:2605.27879v1 Announce Type: new Abstract: Explainable AI (XAI) helps users interpret model behavior and identify potential faults. Agentic XAI systems use Large Language Models (LLMs) to make ex

model-releasesarxiv-cs-ai
28 May 2026
Safety

TRACES: Proactive Safety Auditing for Multi-Turn LLM Agents via Trajectory-State Modeling

DGX agent

arXiv:2605.27690v1 Announce Type: new Abstract: LLM agents increasingly operate through multi-turn tool use and environment interaction, where safety risks often emerge from intermediate steps long be

safetyarxiv-cs-cl
28 May 2026
Safety

ChainCaps: Composition-Safe Tool-Using Agents via Monotonic Capability Attenuation

DGX agent

arXiv:2605.26542v1 Announce Type: cross Abstract: Tool-using agents increasingly operate in open-ended deployment environments, where they compose file systems, web APIs, code interpreters, and enterp

safetyarxiv-cs-ai
27 May 2026
Safety

FinHarness: An Inline Lifecycle Safety Harness for Finance LLM Agents

DGX agent

arXiv:2605.27333v1 Announce Type: new Abstract: Finance LLM agents must simultaneously block prompt-induced unauthorized actions and approve legitimate multi-step business workflows. However, boundary

safetyarxiv-cs-cl
27 May 2026
Agents

GENESIS: Harnessing AI Agents for Autonomous 6G RAN Synthesis, Research, and Testing

DGX agent

arXiv:2605.27360v1 Announce Type: cross Abstract: Cellular research and development (R&D) is throttled by six structural processes that each consume months of manual engineering work per iteration: (i

agentsarxiv-cs-ai
27 May 2026
Model Releases

MobileExplorer: Accelerating On-Device Inference for Mobile GUI Agents via Online Exploration

DGX agent

arXiv:2605.26546v1 Announce Type: new Abstract: Mobile graphical user interface (GUI) agents enable AI models to autonomously operate smartphones on behalf of users. However, most existing systems foc

model-releasesarxiv-cs-ai
27 May 2026
Agents

SetupX: Can LLM Agents Learn from Past Failures in Functionality-Correct Code Repository Setup?

DGX agent

arXiv:2605.26186v1 Announce Type: cross Abstract: Functionality-correct repository setup aims to configure execution environments (e.g., dependencies, build scripts) to successfully execute a reposito

agentsarxiv-cs-ai
27 May 2026
Applications

UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action

DGX agent

arXiv:2510.17790v3 Announce Type: replace-cross Abstract: Computer-use agents face a fundamental limitation. They rely exclusively on primitive GUI actions (click, type, scroll), creating brittle exec

applicationsarxiv-cs-cl
27 May 2026
Model Releases

Anticipate and Learn: Unleashing Idle-Time Compute in Proactive Agents

DGX agent

arXiv:2605.25971v1 Announce Type: new Abstract: While AI agents demonstrate remarkable capabilities in reasoning and tool use, they remain fundamentally reactive: they compute responses only after exp

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows

DGX agent

arXiv:2605.24219v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous agents that reason, use tools, and act over multiple steps. Yet most hallucination

model-releasesarxiv-cs-ai
26 May 2026
Agents

Decoding ML Decision: An Agentic Reasoning Framework for Large-Scale Ranking System

DGX agent

arXiv:2602.18640v2 Announce Type: replace Abstract: Modern large-scale ranking systems operate within a sophisticated landscape of competing objectives, operational constraints, and evolving product r

agentsarxiv-cs-ai
26 May 2026
Agents

Detectify debuts MCP server to let AI agents find and fix vulnerabilities in real time

DGX agent

Application security platform company Detectify AB today launched the Detectify MCP Server, a new integration layer that plugs the company’s security testing engines into artificial intelligence-drive

agentssiliconangle
26 May 2026
Agents

EvoSci: A Bio-Inspired Multi-Agent Framework for the Evolution of Scientific Discovery

DGX agent

arXiv:2605.24018v1 Announce Type: new Abstract: Large language models (LLMs), have shown strong potential in scientific discovery, yet existing methods still face substantial challenges in the design

agentsarxiv-cs-ai
26 May 2026
Model Releases

GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning

DGX agent

arXiv:2605.25200v1 Announce Type: new Abstract: Travel planning is a realistic task for evaluating the planning and tool-use abilities of LLM agents. However, existing benchmarks typically assume only

model-releasesarxiv-cs-cl
26 May 2026
Local Ai

Hera: Learning Long-Horizon Coordination for Device-Cloud Collaborative LLM Agents

DGX agent

arXiv:2605.24598v1 Announce Type: new Abstract: Large language model (LLM) agents excel at solving complex long-horizon tasks through autonomous interaction with environments. However, their real-worl

local-aiarxiv-cs-ai
26 May 2026
Model Releases

IterInject: Indirect Prompt Injection Against LLM Agents via Feedback-Guided Iterative Optimization

DGX agent

arXiv:2605.24659v1 Announce Type: new Abstract: LLM-based agents are increasingly deployed for complex tasks requiring planning, tool use, and interaction with external services. Their reliance on unt

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries

DGX agent

arXiv:2508.15760v2 Announce Type: replace-cross Abstract: Tool calling has emerged as a critical capability for AI agents. In contrast to conventional tool calling frameworks that rely on static, prov

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

MDIA: A Multi-Agent Diagnostic Intelligence Pipeline on HealthBench Professional

DGX agent

arXiv:2605.24699v1 Announce Type: new Abstract: Most reported gains on agentic-LLM clinical benchmarks are often attributed to prompt engineering, yet our results suggest that larger improvements can

model-releasesarxiv-cs-ai
26 May 2026
Safety

Multi-Agent Coordination Adaptation via Structure-Guided Orchestration

DGX agent

arXiv:2605.25746v1 Announce Type: cross Abstract: As large language model (LLM)-based multi-agent systems scale to handle increasingly complex tasks, balancing structural stability and dynamic adaptab

safetyarxiv-cs-ai
26 May 2026
Agents

Multi-Agent Specification-based Metamorphic Testing of FMU-Based Simulations

DGX agent

arXiv:2605.25101v1 Announce Type: cross Abstract: In many industrial domains, the Functional Mock-up Interface (FMI) is used to exchange simulation models as Functional Mock-up Units (FMUs) across dif

agentsarxiv-cs-ai
26 May 2026
Safety

Toward Reliable Design of LLM-Enabled Agentic Workflows: Optimizing Latency-Reliability-Cost Tradeoffs

DGX agent

arXiv:2605.23929v1 Announce Type: new Abstract: Modern AI systems increasingly rely on workflows composed of multiple interacting agents, some powered by large language models (LLMs) and others by con

safetyarxiv-cs-ai
26 May 2026
Safety

When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs

DGX agent

arXiv:2605.24202v1 Announce Type: new Abstract: Multi-agent LLM workflows route inference through specialized roles to lift end-task accuracy, but jointly training those roles with reinforcement learn

safetyarxiv-cs-ai
26 May 2026
Model Releases

Agentic Proving for Program Verification

DGX agent

arXiv:2605.23772v1 Announce Type: new Abstract: Agentic systems have recently emerged as state-of-the-art approaches for automated theorem proving in formal mathematics. To assess how far these capabi

model-releasesarxiv-cs-ai
25 May 2026
Safety

Foundation Protocol: A Coordination Layer for Agentic Society

DGX agent

arXiv:2605.23218v1 Announce Type: new Abstract: Autonomous agents are moving from tools into a layer of social infrastructure: they browse, purchase, deploy software, manage systems, and increasingly

safetyarxiv-cs-ai
25 May 2026
← Previous
1…119120121122123…375
Next →