AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,289 results
Model Releases

Continual Harness: Online Adaptation for Self-Improving Foundation Agents

DGX agent

arXiv:2605.09998v1 Announce Type: cross Abstract: Coding harnesses such as Claude Code and OpenHands wrap foundation models with tools, memory, and planning, but no equivalent exists for embodied agen

model-releasesarxiv-cs-ai
12 May 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

CrackMeBench: Binary Reverse Engineering for Agents

DGX agent

arXiv:2605.10597v1 Announce Type: cross Abstract: Benchmarks for coding agents increasingly measure source-level software repair, and cybersecurity benchmarks increasingly measure broad capture-the-fl

model-releasesarxiv-cs-ai
12 May 2026
Safety

Debugging the Debuggers: Failure-Anchored Structured Recovery for Software Engineering Agents

DGX agent

arXiv:2605.08717v1 Announce Type: cross Abstract: Software engineering agents are increasingly deployed in evaluable engineering environments, yet post-failure recovery remains costly, manual, and ad

safetyarxiv-cs-ai
12 May 2026
Model Releases

DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents

DGX agent

arXiv:2605.09679v1 Announce Type: cross Abstract: Medical vision-language models (VLMs) and AI agents have made significant progress in learning to analyze and reason about clinical images. However, e

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

EmbodiSkill: Skill-Aware Reflection for Self-Evolving Embodied Agents

DGX agent

arXiv:2605.10332v1 Announce Type: new Abstract: Embodied agents can benefit from skills that guide object search, action execution, and state changes across diverse environments. Since embodied enviro

model-releasesarxiv-cs-ai
12 May 2026
Agents

Generating synthetic electronic health record data using agent-based models to evaluate machine learning robustness under mass casualty incidents

DGX agent

arXiv:2605.09951v1 Announce Type: new Abstract: ML models in healthcare are typically evaluated using curated real-world EHR data. A key limitation of such evaluations is that they may fail to assess

agentsarxiv-cs-lg
12 May 2026
Agents

HAMLET: A Hierarchical and Adaptive Multi-Agent Framework for Live Embodied Theatrics

DGX agent

arXiv:2507.15518v5 Announce Type: replace Abstract: Creating an immersive and interactive theatrical experience is a long-term goal in the field of interactive narrative. The emergence of large langua

agentsarxiv-cs-ai
12 May 2026
Research

How Mobile World Model Guides GUI Agents?

DGX agent

arXiv:2605.10347v1 Announce Type: new Abstract: Recent advances in vision-language models have enabled mobile GUI agents to perceive visual interfaces and execute user instructions, but reliable predi

researcharxiv-cs-ai
12 May 2026
Safety

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization

DGX agent

arXiv:2605.08978v1 Announce Type: new Abstract: Recent advancements in agentic test-time scaling allow models to gather environmental feedback before committing to final actions. A key limitation of e

safetyarxiv-cs-ai
12 May 2026
Model Releases

Log analysis is necessary for credible evaluation of AI agents

DGX agent

arXiv:2605.08545v1 Announce Type: new Abstract: Agent benchmarks typically report only final outcomes: pass or fail. This threatens evaluation credibility in three ways. First, scores may be inflated

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MDGYM: Benchmarking AI Agents on Molecular Simulations

DGX agent

arXiv:2605.08941v1 Announce Type: new Abstract: The promise of AI-driven scientific discovery hinges on whether AI agents can autonomously design and execute the computational workflows that underpin

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Nautilus Compass: Black-box Persona Drift Detection for Production LLM Agents

DGX agent

arXiv:2605.09863v1 Announce Type: cross Abstract: Production LLM coding agents drift over long sessions: they forget user-specified constraints, slip into mistakes the user already flagged, and confab

model-releasesarxiv-cs-ai
12 May 2026
Hardware

Octopus Protocol: One-Shot Hardware Discovery and Control for AI Agents via Infrastructure-as-Prompts

DGX agent

arXiv:2605.09055v1 Announce Type: cross Abstract: Recent agentic-robotics systems, from Code-asPolicies to modern vision-language-action (VLA) foundation models, presuppose that drivers, SDKs, or ROS-

hardwarearxiv-cs-ai
12 May 2026
Model Releases

Oracle Poisoning: Corrupting Knowledge Graphs to Weaponise AI Agent Reasoning

DGX agent

arXiv:2605.09822v1 Announce Type: cross Abstract: We define Oracle Poisoning, an attack class in which an adversary corrupts a structured knowledge graph that AI agents query at runtime via tool-use p

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM Agents

DGX agent

arXiv:2605.08876v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous agents that execute tool-augmented, multi-step tasks, where latency is a critical f

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage

DGX agent

arXiv:2604.01527v3 Announce Type: replace-cross Abstract: Production deployment of AI coding agents requires fast, reproducible evaluation signals. Existing industrial practices trade off speed and fi

model-releasesarxiv-cs-ai
12 May 2026
Safety

Route by State, Recover from Trace: STAR with Failure-Aware Markov Routing for Multi-Agent Spatiotemporal Reasoning

DGX agent

arXiv:2605.10057v1 Announce Type: new Abstract: Compositional spatiotemporal reasoning often requires a system to invoke multiple heterogeneous specialists, such as geometric, temporal, topological, a

safetyarxiv-cs-ai
12 May 2026
Agents

ScholarPeer: A Context-Aware Multi-Agent Framework for Automated Peer Review

DGX agent

arXiv:2601.22638v2 Announce Type: replace-cross Abstract: The exponential growth of machine learning submissions has strained the traditional peer review process, resulting in slow feedback loops for

agentsarxiv-cs-ai
12 May 2026
Agents

Simulus: Combining Improvements in Sample-Efficient World Model Agents

DGX agent

arXiv:2502.11537v4 Announce Type: replace-cross Abstract: World models (WMs) represent the frontier of sample-efficient reinforcement learning, but their complexity leaves many promising improvements

agentsarxiv-cs-ai
12 May 2026
Safety

The Bystander Effect in Multi-Agent Reasoning: Quantifying Cognitive Loafing in Collaborative Interactions

DGX agent

arXiv:2605.10698v1 Announce Type: cross Abstract: Multi-agent systems (MAS) assume that collaborating inherently improves Large Language Model (LLM) reasoning. We challenge this by demonstrating that

safetyarxiv-cs-ai
12 May 2026
Model Releases

The Trap of Trajectory: Towards Understanding and Mitigating Spurious Correlations in Agentic Memory

DGX agent

arXiv:2605.09330v1 Announce Type: cross Abstract: Agentic memory enables LLMs to persist information beyond a single context window and reuse it in later decisions, but it also introduces a new vulner

model-releasesarxiv-cs-ai
12 May 2026
Agents

When Agents Say One Thing and Do Another: Validating Elicited Beliefs from LLMs

DGX agent

arXiv:2602.06286v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in high-stakes settings where good decisions require forming beliefs over the probability of

agentsarxiv-cs-ai
12 May 2026
Agents

AGILE: Hand-Object Interaction Reconstruction from Video via Agentic Generation

DGX agent

arXiv:2602.04672v3 Announce Type: replace Abstract: Reconstructing dynamic hand-object interactions from monocular videos is critical for dexterous manipulation data collection and creating realistic

agentsarxiv-cs-cv
11 May 2026
Model Releases

Ask Early, Ask Late, Ask Right: When Does Clarification Timing Matter for Long-Horizon Agents?

DGX agent

arXiv:2605.07937v1 Announce Type: new Abstract: Long-horizon AI agents execute complex workflows spanning hundreds of sequential actions, yet a single wrong assumption early on can cascade into irreve

model-releasesarxiv-cs-cl
11 May 2026
Agents

ATHENA: Agentic Team for Hierarchical Evolutionary Numerical Algorithms

DGX agent

arXiv:2512.03476v2 Announce Type: replace-cross Abstract: Bridging the gap between theoretical conceptualization and computational implementation is a major bottleneck in Scientific Computing (SciC) a

agentsarxiv-cs-ai
11 May 2026
Local Ai

EditRefiner: A Human-Aligned Agentic Framework for Image Editing Refinement

DGX agent

arXiv:2605.07457v1 Announce Type: new Abstract: Recent text-guided image editing (TIE) models have made remarkable progress, yet edited images still frequently suffer from fine-grained issues such as

local-aiarxiv-cs-cv
11 May 2026
Safety

From Specification to Deployment: Empirical Evidence from a W3C VC + DID Trust Infrastructure for Autonomous Agents

DGX agent

arXiv:2605.06738v1 Announce Type: cross Abstract: Autonomous AI agents now transact at production scale -- 69,000 bots executing 165 million transactions across 50 million USDC in cumulative volume on

safetyarxiv-cs-ai
11 May 2026
Model Releases

HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents

DGX agent

arXiv:2605.07177v1 Announce Type: cross Abstract: Existing multimodal search agents process target entities sequentially, issuing one tool call per entity and accumulating redundant interaction rounds

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

InterLV-Search: Benchmarking Interleaved Multimodal Agentic Search

DGX agent

arXiv:2605.07510v1 Announce Type: cross Abstract: Existing benchmarks for multimodal agentic search evaluate multimodal search and visual browsing, but visual evidence is either confined to the input

model-releasesarxiv-cs-cl
11 May 2026
Safety

Multi-Objective Multi-Agent Bandits: From Learning Efficiency to Fairness Optimization

DGX agent

arXiv:2605.06864v1 Announce Type: new Abstract: We study multi-objective multi-agent multi-armed bandits (MO-MA-MAB) under stochastic rewards, where agents observe heterogeneous reward vectors and com

safetyarxiv-cs-lg
11 May 2026
Safety

OrchJail: Jailbreaking Tool-Calling Text-to-Image Agents by Orchestration-Guided Fuzzing

DGX agent

arXiv:2605.07414v1 Announce Type: cross Abstract: Tool-calling text-to-image (T2I) agents can plan and execute multi-step tool chains to accomplish complex generation and editing queries. However, thi

safetyarxiv-cs-ai
11 May 2026
Model Releases

Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents

DGX agent

arXiv:2605.07630v1 Announce Type: cross Abstract: When a phone-use agent avoids harm, does that show safety, or simply inability to act? Existing evaluations often cannot tell. A harmful outcome may b

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair

DGX agent

arXiv:2605.07001v1 Announce Type: cross Abstract: Architectural code smells erode software maintainability and are costly to repair manually, yet unlike localized bugs, they require cross-module reaso

model-releasesarxiv-cs-cl
11 May 2026
Agents

Towards Autonomous Business Intelligence via Data-to-Insight Discovery Agent

DGX agent

arXiv:2605.07202v1 Announce Type: new Abstract: Transforming fragmented enterprise data into actionable insights remains a significant challenge for LLMs, constrained by complex database schemas, limi

agentsarxiv-cs-ai
11 May 2026
Agents

ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration

DGX agent

arXiv:2605.03042v1 Announce Type: cross Abstract: This report describes ARIS (Auto-Research-in-sleep), an open-source research harness for autonomous research, including its architecture, assurance me

agentsarxiv-cs-ai
7 May 2026
Agents

Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning

DGX agent

arXiv:2605.04304v1 Announce Type: cross Abstract: Advanced chart question answering requires both precise perception of small visual elements and multi-step reasoning across several subplots. While ex

agentsarxiv-cs-cl
7 May 2026
Agents

Meta-Learning and Meta-Reinforcement Learning -- Tracing the Path towards DeepMind's Adaptive Agent

DGX agent

arXiv:2602.19837v3 Announce Type: replace-cross Abstract: Humans are highly effective at utilizing prior knowledge to adapt to novel tasks, a capability that standard machine learning models struggle

agentsarxiv-cs-lg
7 May 2026
Model Releases

MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents

DGX agent

arXiv:2605.03952v1 Announce Type: cross Abstract: Coding agents often pass per-prompt safety review yet ship exploitable code when their tasks are decomposed into routine engineering tickets. The chal

model-releasesarxiv-cs-ai
7 May 2026
Safety

Quantifying Trust: Financial Risk Management for Trustworthy AI Agents

DGX agent

arXiv:2604.03976v2 Announce Type: replace Abstract: Prior work on trustworthy AI emphasizes model-internal properties such as bias mitigation, adversarial robustness, and interpretability. As AI syste

safetyarxiv-cs-ai
7 May 2026
Agents

SymptomAI: Towards a Conversational AI Agent for Everyday Symptom Assessment

DGX agent

arXiv:2605.04012v1 Announce Type: new Abstract: Language models excel at diagnostic assessments on currated medical case-studies and vignettes, performing on par with, or better than, clinical profess

agentsarxiv-cs-ai
7 May 2026
Local Ai

The Hive Mind is a Single Reinforcement Learning Agent

DGX agent

arXiv:2410.17517v5 Announce Type: replace-cross Abstract: Decision-making is an essential attribute of any intelligent agent or group. Natural systems are known to converge to effective strategies thr

local-aiarxiv-cs-ai
7 May 2026
Agents

A Compound AI Agent for Conversational Grant Discovery

DGX agent

arXiv:2605.02366v1 Announce Type: new Abstract: Research funding discovery remains fundamentally fragmented: researchers navigate disparate agency portals (e.g., in the United States, NSF, NIH, DARPA,

agentsarxiv-cs-ai
6 May 2026
Agents

AI-Generated Smells: An Analysis of Code and Architecture in LLM and Agent-Driven Development

DGX agent

arXiv:2605.02741v1 Announce Type: cross Abstract: The promise of Large Language Models in automated software engineering is often measured by functional correctness, overlooking the critical issue of

agentsarxiv-cs-ai
6 May 2026
Safety

Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents

DGX agent

arXiv:2604.06132v2 Announce Type: replace Abstract: Large language models are increasingly deployed as autonomous agents for multi-step workflows in real-world software environments. However, existing

safetyarxiv-cs-ai
6 May 2026
Model Releases

ContextCov: Deriving and Enforcing Executable Constraints from Agent Instruction Files

DGX agent

arXiv:2603.00822v2 Announce Type: replace-cross Abstract: As Large Language Model (LLM) agents increasingly execute complex, autonomous software engineering tasks, developers rely on natural language

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Coordination as an Architectural Layer for LLM-Based Multi-Agent Systems

DGX agent

arXiv:2605.03310v1 Announce Type: cross Abstract: Multi-agent LLM systems fail in production at rates between 41% and 87%, mostly due to coordination defects rather than base-model capability. Existin

model-releasesarxiv-cs-lg
6 May 2026
Model Releases

DataClaw: A Process-Oriented Agent Benchmark for Exploratory Real-World Data Analysis

DGX agent

arXiv:2605.02503v1 Announce Type: new Abstract: Evaluating autonomous data analysis agents requires testing their ability to perform exploratory analysis in underexplored data environments. However, m

model-releasesarxiv-cs-ai
6 May 2026
Agents

EngiAgent: Fully Connected Coordination of LLM Agents for Solving Open-ended Engineering Problems with Feasible Solutions

DGX agent

arXiv:2605.02289v1 Announce Type: new Abstract: Engineering problem solving is central to real-world decision-making, requiring mathematical formulations that not only represent complex problems but a

agentsarxiv-cs-ai
6 May 2026
← Previous
1…7172737475…236
Next →