AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,153 results
Safety

BiCAA: Bidirectional Credit Assignment for Search-Augmented Agent

DGX agent

arXiv:2608.01321v1 Announce Type: new Abstract: Multi-step search is a fundamental capability for search agents, enabling them to iteratively acquire, refine, and integrate external evidence for compl

safetyarxiv-cs-cl
4 Aug 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Global Optimization and Inference-Time Region Grafting for Agentic Workflows

DGX agent

arXiv:2608.02353v1 Announce Type: new Abstract: Recent advances in agentic workflow optimization automate workflow design through task-specific workflow search or input-conditioned architecture select

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

HopRefusalBench: Diagnosing Refusal Failures in Search-Augmented Agents for Multi-Hop Reasoning

DGX agent

arXiv:2608.01358v1 Announce Type: new Abstract: Search-augmented large language model agents are increasingly capable of solving knowledge-intensive tasks, but their behavior when a multi-hop question

model-releasesarxiv-cs-cl
4 Aug 2026
Agents

Self-Supervised Skill Optimization

DGX agent

arXiv:2607.28777v1 Announce Type: new Abstract: Agent skills provide frozen large language model (LLM) agents with reusable procedural guidance, and recent work shows that such skills can be optimized

agentsarxiv-cs-cl
3 Aug 2026
Agents

Transcript-Managed Transformers: Monotone Multi-Agent Collapse and Universality with Two Pop-Enabled Transcripts

DGX agent

arXiv:2607.29496v1 Announce Type: new Abstract: We study transcript management for fixed, finite-precision causal Transformers. A transcript is partitioned into channels of bounded blocks. Each transi

agentsarxiv-cs-lg
3 Aug 2026
Model Releases

LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation

DGX agent

arXiv:2607.27353v1 Announce Type: new Abstract: Agentic retrieval-augmented generation systems can produce answers that appear grounded while failing at the evidence, tool-contract, authorization, or

model-releasesarxiv-cs-cl
31 Jul 2026
Agents

Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork

DGX agent

arXiv:2607.27177v1 Announce Type: new Abstract: Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents. Most current ad-hoc teamwork (AHT) approaches assume t

agentsarxiv-cs-ai
31 Jul 2026
Safety

Procedural Fairness in Multi-Agent Bandits

DGX agent

arXiv:2601.10600v2 Announce Type: replace-cross Abstract: In the context of multi-agent multi-armed bandits (MA-MAB), fairness is often reduced to outcomes: maximizing welfare, reducing inequality, or

safetyarxiv-cs-lg
31 Jul 2026
Model Releases

Safety-Gated Agentic Supervisory Control on a Coupled Distillation Benchmark: Regime Map, Auditable Gate, and Co-Design Findings

DGX agent

arXiv:2607.27849v1 Announce Type: cross Abstract: An open-weight LLM can write composition setpoints every five minutes. What a plant still needs is a hard check: named constraints, logged margins, an

model-releasesarxiv-cs-lg
31 Jul 2026
Safety

SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution

DGX agent

arXiv:2607.26784v1 Announce Type: new Abstract: Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet standard agentic reinforcement learnin

safetyarxiv-cs-lg
30 Jul 2026
Safety

WikiLoop: Jointly Learning to Build and Navigate Agent-Native Wikis with Downstream Feedback

DGX agent

arXiv:2607.26604v1 Announce Type: new Abstract: Knowledge-base construction and querying are typically optimized in isolation: retrieval-augmented agents operate over a fixed, externally maintained in

safetyarxiv-cs-cl
30 Jul 2026
Agents

Cardiologent: Multi-Agent Clinical Decision Support for Patient-Level Arrhythmia Assessment, Urgency, and Management

DGX agent

arXiv:2607.25340v1 Announce Type: new Abstract: The same episode of atrial fibrillation is a minor finding in a healthy adult and grounds for anticoagulation in an elderly patient with hypertension: i

agentsarxiv-cs-ai
29 Jul 2026
Safety

ContractHIL-HLS: Contract-Aligned Multi-Agent Workflow with Hardware-in-the-Loop Feedback for HLS Design

DGX agent

arXiv:2607.25283v1 Announce Type: new Abstract: This paper presents ContractHIL-HLS, a contract-aligned multi-agent workflow for practical high-level synthesis (HLS) engineering. The workflow makes th

safetyarxiv-cs-ai
29 Jul 2026
Model Releases

COVENANT: Natural-Language Workflow Compilation for Aligned Agent Execution

DGX agent

arXiv:2607.25400v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly entrusted with natural-language workflow instructions (e.g., retail-payment policies) that specify no

model-releasesarxiv-cs-ai
29 Jul 2026
Agents

From Naive RAG to Deep Agentic Retrieval: An Evolving Context Engineering Pipeline for Regulatory Compliance

DGX agent

arXiv:2607.24791v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) is the dominant paradigm for applying large language models (LLMs) to enterprise document corpora, yet naive impl

agentsarxiv-cs-ai
29 Jul 2026
Model Releases

Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification

DGX agent

arXiv:2607.25904v1 Announce Type: new Abstract: Graphical user interface task evaluation aims to determine whether a GUI agent has successfully completed a user instruction. Automated GUI task evaluat

model-releasesarxiv-cs-ai
29 Jul 2026
Model Releases

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation

DGX agent

arXiv:2607.25891v1 Announce Type: new Abstract: Evaluating AI agents in interactive environments is hindered by fragmented tasks, scaffolds, verifiers, and scoring rules. Existing efforts focus on nar

model-releasesarxiv-cs-ai
29 Jul 2026
Safety

SearchArt: Training Long-Horizon Search Agent with Scalable Synthetic and Verified Task

DGX agent

arXiv:2607.24850v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have enabled search agents to autonomously tackle complex tasks across extended search and reasoning h

safetyarxiv-cs-lg
29 Jul 2026
Model Releases

WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing

DGX agent

arXiv:2607.25765v1 Announce Type: new Abstract: Enterprise agents often need to integrate heterogeneous knowledge sources: documents for narrative facts, tables for computation, and dependency graphs

model-releasesarxiv-cs-cl
29 Jul 2026
Model Releases

AlloBench: Measuring Online Tool Allocation Capability in LLM Agents

DGX agent

arXiv:2607.23332v1 Announce Type: new Abstract: Creating a reusable tool is an investment: an agent pays a fixed cost now in exchange for the potential of future reuse. Therefore, a user should prefer

model-releasesarxiv-cs-lg
28 Jul 2026
Model Releases

ConsistencyGate: Preventing Memory Contamination in LLM Agents via Self-Consistency Admission Control

DGX agent

arXiv:2607.22962v1 Announce Type: new Abstract: LLM agents that operate over many turns accumulate facts in an external memory store and reuse them as premises for downstream reasoning. A hallucinated

model-releasesarxiv-cs-ai
28 Jul 2026
Safety

CoopReflect: Towards Natural Language Communication for Cooperative Autonomous Driving via Multi-Agent Learning

DGX agent

arXiv:2505.18334v2 Announce Type: replace-cross Abstract: Past work has demonstrated that autonomous vehicles can drive more safely if they communicate with each other. However, this communication is

safetyarxiv-cs-ai
28 Jul 2026
Agents

Knowledge-Centric Agents for Workflow Generation in ComfyUI

DGX agent

arXiv:2607.15845v2 Announce Type: replace Abstract: Workflow generation in visual creation systems such as ComfyUI demands not only syntactic accuracy but also expert-level reasoning over modular comp

agentsarxiv-cs-ai
28 Jul 2026
Local Ai

Let AI Agents Translate Networks, Not Reason About Them

DGX agent

arXiv:2607.22947v1 Announce Type: new Abstract: A formal model enables verifying reachability, localizing an outage, or anticipating the blast radius of a change. Yet, virtually no production network

local-aiarxiv-cs-ai
28 Jul 2026
Model Releases

MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents

DGX agent

arXiv:2607.23870v1 Announce Type: cross Abstract: Smart-city airspace is transforming Uncrewed Aerial Vehicles (UAVs) from passive sensing platforms into cyber-physical decision makers that must follo

model-releasesarxiv-cs-ai
28 Jul 2026
Agents

Multi-Agent Privacy Game in Federated Learning: A Unified Mean-Field View

DGX agent

arXiv:2607.23029v1 Announce Type: cross Abstract: Federated learning enables collaborative model training across distributed clients without centralising their data, yet privacy remains a persistent c

agentsarxiv-cs-ai
28 Jul 2026
Safety

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation

DGX agent

arXiv:2607.24720v1 Announce Type: cross Abstract: Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear. Existing models are tra

safetyarxiv-cs-ai
28 Jul 2026
Model Releases

UAV-ON: A Benchmark for Open-World Object Goal Navigation with Aerial Agents

DGX agent

arXiv:2508.00288v5 Announce Type: replace-cross Abstract: Aerial navigation is a fundamental yet underexplored capability in embodied intelligence, enabling agents to operate in large-scale, unstructu

model-releasesarxiv-cs-cv
28 Jul 2026
Safety

What Can Be Enforced? A Theory of Certified Runtime Safety for Tool-Using Agents

DGX agent

arXiv:2607.22868v1 Announce Type: new Abstract: Runtime guardrails act before irreversible tool calls, but their guarantees depend on what policy state is representable, what a judge observes, and whe

safetyarxiv-cs-ai
28 Jul 2026
Model Releases

AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use

DGX agent

arXiv:2607.20536v1 Announce Type: new Abstract: Tool-use agents that address day-to-day digital tasks such as ordering groceries must not only operate applications, but also interact with the user, e.

model-releasesarxiv-cs-ai
24 Jul 2026
Safety

EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization

DGX agent

arXiv:2607.21013v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have achieved impressive performance in multimodal emotion recognition (MER) tasks and lifted MER to a new leve

safetyarxiv-cs-ai
24 Jul 2026
Model Releases

Frontier Financial Judgement: Can agents tell what might move a stock?

DGX agent

arXiv:2607.20645v1 Announce Type: cross Abstract: We introduce Frontier Financial Judgement, a challenging new benchmark developed in collaboration with professional equity analysts to assess agents'

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

pAI-Econ-claude: A Gated Human-in-the-Loop Multi-Agent Architecture for AI-Assisted Economic Theory Development

DGX agent

arXiv:2607.21268v1 Announce Type: cross Abstract: In many social-science research tasks, such as economics, LLM-based agents must produce outputs for which no cheap, task-complete, machine-readable co

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

PersonaTrail: Benchmarking Personalized Web Agents through Browsing Trails

DGX agent

arXiv:2607.20482v1 Announce Type: new Abstract: Recent advances in large language models have enabled web agents to autonomously execute complex tasks. In practice, users frequently provide underspeci

model-releasesarxiv-cs-ai
24 Jul 2026
Agents

TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework

DGX agent

arXiv:2511.05385v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) utilizes external knowledge to augment Large Language Models' (LLMs) reliability. For flexibility, agenti

agentsarxiv-cs-ai
24 Jul 2026
Model Releases

DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations

DGX agent

arXiv:2607.19865v1 Announce Type: new Abstract: As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become critical for enabling general-purpose

model-releasesarxiv-cs-ai
23 Jul 2026
Safety

Harnessing Disagreement: Detecting Correlated Agreement Blindness in Multi-Agent Triage

DGX agent

arXiv:2607.19899v1 Announce Type: cross Abstract: Disagreement-triggered escalation can create a structural blind spot in multi-agent arbitration: as base learners improve, they tend to converge, weak

safetyarxiv-cs-lg
23 Jul 2026
Agents

Silent Failures in Multimodal Agentic Search:A Diagnostic Taxonomy and Cross-Judge Evaluation

DGX agent

arXiv:2607.19793v1 Announce Type: new Abstract: Multimodal agentic search systems increasingly rely on external tools to answer knowledge-intensive visual questions. However, existing evaluations main

agentsarxiv-cs-ai
23 Jul 2026
Model Releases

A Self-Evolving Agent for Longitudinal Personal Health Management

DGX agent

arXiv:2607.13940v1 Announce Type: new Abstract: Personal health management unfolds over repeated encounters, yet most health AI systems treat each request in isolation. We developed HealthClaw, an ope

model-releasesarxiv-cs-ai
16 Jul 2026
Safety

AI-Native Insurance for Agentic AI: Pricing, Underwriting, and End-to-End Automation

DGX agent

arXiv:2607.13230v1 Announce Type: new Abstract: Agentic AI introduces new insurance challenges because autonomous AI systems can make decisions, invoke tools, modify external environments, and interac

safetyarxiv-cs-ai
16 Jul 2026
Model Releases

DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments

DGX agent

arXiv:2607.13465v1 Announce Type: cross Abstract: LLM-based agents have rapidly improved at operating individual digital environments such as mobile applications, desktop systems, and smart homes. How

model-releasesarxiv-cs-ai
16 Jul 2026
Safety

Explaining Reinforcement Learning Agents via Inductive Logic Programming

DGX agent

arXiv:2607.13655v1 Announce Type: new Abstract: Explainable Reinforcement Learning (XRL) seeks to make Reinforcement Learning (RL) policies more transparent and interpretable, a key requirement in saf

safetyarxiv-cs-ai
16 Jul 2026
Safety

Learning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies

DGX agent

arXiv:2604.00830v3 Announce Type: replace-cross Abstract: Test-Time Learning (TTL) enables language agents to iteratively refine their performance through repeated interactions with the environment at

safetyarxiv-cs-ai
16 Jul 2026
Model Releases

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory

DGX agent

arXiv:2607.10350v1 Announce Type: cross Abstract: Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general runtime la

model-releasesarxiv-cs-ro
15 Jul 2026
Model Releases

Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric

DGX agent

arXiv:2607.12469v1 Announce Type: cross Abstract: Many agent-safety evaluation results are not yet load-bearing evidence: identical nominal outcomes (task success, attack success, monitor scores) may

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Agentic systems for breast cancer treatment recommendations

DGX agent

arXiv:2607.12051v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly being explored for clinical decision support, but their reliability in complex oncology treatment planning

model-releasesarxiv-cs-cl
15 Jul 2026
Agents

EvoGraph-R1: Self-Evolving Multimodal Knowledge Hypergraphs for Agentic Retrieval

DGX agent

arXiv:2607.12764v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) has emerged as a critical paradigm for grounding Multimodal Large Language Models (MLLMs) in external knowledge. Re

agentsarxiv-cs-cv
15 Jul 2026
Model Releases

Rethinking the Evaluation of Harness Evolution for Agents

DGX agent

arXiv:2607.12227v1 Announce Type: new Abstract: We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit test cases to search for harness co

model-releasesarxiv-cs-ai
15 Jul 2026
← Previous
1…5455565758…233
Next →