AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,289 results
Local Ai

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning

DGX agent

arXiv:2601.15724v2 Announce Type: replace Abstract: Long-form video understanding remains a fundamental challenge for current Video Large Language Models. Most existing models rely on static reasoning

local-aiarxiv-cs-cv
21 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Agents

What Makes AI Research Replicable? Executable Knowledge Graphs as Scientific Knowledge Representations

DGX agent

arXiv:2510.17795v3 Announce Type: replace Abstract: Replicating AI research is a crucial yet challenging task for large language model (LLM) agents. Existing approaches often struggle to generate exec

agentsarxiv-cs-cl
21 Apr 2026
Model Releases

ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence

DGX agent

arXiv:2603.24621v2 Announce Type: replace Abstract: We introduce ARC-AGI-3, an interactive benchmark for studying agentic intelligence through novel, abstract, turn-based environments in which agents

model-releasesarxiv-cs-ai
20 Apr 2026
Model Releases

ChemGraph-XANES: An Agentic Framework for XANES Simulation and Analysis

DGX agent

arXiv:2604.16205v1 Announce Type: cross Abstract: Computational X-ray absorption near-edge structure (XANES) is widely used to probe local coordination environments, oxidation states, and electronic s

model-releasesarxiv-cs-ai
20 Apr 2026
Model Releases

HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?

DGX agent

arXiv:2604.15415v1 Announce Type: cross Abstract: Large language models (LLMs) have evolved into autonomous agents that rely on open skill ecosystems (e.g., ClawHub and Skills.Rest), hosting numerous

model-releasesarxiv-cs-ai
20 Apr 2026
Model Releases

Preference Estimation via Opponent Modeling in Multi-Agent Negotiation

DGX agent

arXiv:2604.15687v1 Announce Type: new Abstract: Automated negotiation in complex, multi-party and multi-issue settings critically depends on accurate opponent modeling. However, conventional numerical

model-releasesarxiv-cs-cl
20 Apr 2026
Model Releases

RiskWebWorld: A Realistic Interactive Benchmark for GUI Agents in E-commerce Risk Management

DGX agent

arXiv:2604.13531v1 Announce Type: cross Abstract: Graphical User Interface (GUI) agents show strong capabilities for automating web tasks, but existing interactive benchmarks primarily target benign,

model-releasesarxiv-cs-lg
16 Apr 2026
Model Releases

AlphaEval: Evaluating Agents in Production

DGX agent

arXiv:2604.12162v1 Announce Type: new Abstract: The rapid deployment of AI agents in commercial settings has outpaced the development of evaluation methodologies that reflect production realities. Exi

model-releasesarxiv-cs-cl
15 Apr 2026
Model Releases

How memory can affect collective and cooperative behaviors in an LLM-Based Social Particle Swarm

DGX agent

arXiv:2604.12250v1 Announce Type: new Abstract: This study examines how model-specific characteristics of Large Language Model (LLM) agents, including internal alignment, shape the effect of memory on

model-releasesarxiv-cs-ai
15 Apr 2026
Model Releases

SIR-Bench: Evaluating Investigation Depth in Security Incident Response Agents

DGX agent

arXiv:2604.12040v1 Announce Type: cross Abstract: We present SIR-Bench, a benchmark of 794 test cases for evaluating autonomous security incident response agents that distinguishes genuine forensic in

model-releasesarxiv-cs-ai
15 Apr 2026
Model Releases

Transferable Expertise for Autonomous Agents via Real-World Case-Based Learning

DGX agent

arXiv:2604.12717v1 Announce Type: new Abstract: LLM-based autonomous agents perform well on general reasoning tasks but still struggle to reliably use task structure, key constraints, and prior experi

model-releasesarxiv-cs-ai
15 Apr 2026
Model Releases

VULCAN: Vision-Language-Model Enhanced Multi-Agent Cooperative Navigation for Indoor Fire-Disaster Response

DGX agent

arXiv:2604.12831v1 Announce Type: new Abstract: Indoor fire disasters pose severe challenges to autonomous search and rescue due to dense smoke, high temperatures, and dynamically evolving indoor envi

model-releasesarxiv-cs-ro
15 Apr 2026
Agents

When to Forget: A Memory Governance Primitive

DGX agent

arXiv:2604.12007v1 Announce Type: new Abstract: Agent memory systems accumulate experience but currently lack a principled operational metric for memory quality governance -- deciding which memories t

agentsarxiv-cs-ai
15 Apr 2026
Hardware

AIRA_2: Overcoming Bottlenecks in AI Research Agents

DGX agent

arXiv:2603.26499v2 Announce Type: replace Abstract: Existing research has identified three structural performance bottlenecks in AI research agents: (1) synchronous single-GPU execution constrains sam

hardwarearxiv-cs-ai
14 Apr 2026
Model Releases

BankerToolBench: Evaluating AI Agents in End-to-End Investment Banking Workflows

DGX agent

arXiv:2604.11304v1 Announce Type: new Abstract: Existing AI benchmarks lack the fidelity to assess economically meaningful progress on professional workflows. To evaluate frontier AI agents in a high-

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

ClawVM: Harness-Managed Virtual Memory for Stateful Tool-Using LLM Agents

DGX agent

arXiv:2604.10352v1 Announce Type: new Abstract: Stateful tool-using LLM agents treat the context window as working memory, yet today's agent harnesses manage residency and durability as best-effort, c

model-releasesarxiv-cs-ai
14 Apr 2026
Local Ai

CodeComp: Structural KV Cache Compression for Agentic Coding

DGX agent

arXiv:2604.10235v1 Announce Type: new Abstract: Agentic code tasks such as fault localization and patch generation require processing long codebases under tight memory constraints, where the Key-Value

local-aiarxiv-cs-cl
14 Apr 2026
Model Releases

CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification

DGX agent

arXiv:2604.01687v2 Announce Type: replace Abstract: Anthropic proposes the concept of skills for LLM agents to tackle multi-step professional tasks that simple tool invocations cannot address. A tool

model-releasesarxiv-cs-ai
14 Apr 2026
Agents

CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance

DGX agent

arXiv:2508.16644v4 Announce Type: replace Abstract: Diffusion models excel at photorealistic synthesis but struggle with precise object counts, especially in high-density settings. We introduce COUNTL

agentsarxiv-cs-cv
14 Apr 2026
Model Releases

Detecting Safety Violations Across Many Agent Traces

DGX agent

arXiv:2604.11806v1 Announce Type: new Abstract: To identify safety violations, auditors often search over large sets of agent traces. This search is difficult because failures are often rare, complex,

model-releasesarxiv-cs-ai
14 Apr 2026
Safety

EE-MCP: Self-Evolving MCP-GUI Agents via Automated Environment Generation and Experience Learning

DGX agent

arXiv:2604.09815v1 Announce Type: new Abstract: Computer-use agents that combine GUI interaction with structured API calls via the Model Context Protocol (MCP) show promise for automating software tas

safetyarxiv-cs-ai
14 Apr 2026
Model Releases

HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks

DGX agent

arXiv:2604.09937v1 Announce Type: new Abstract: Healthcare administration accounts for over $1 trillion in annual spending, making it a promising target for LLM-based computer-use agents (CUAs). While

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

NetAgentBench: A State-Centric Benchmark for Evaluating Agentic Network Configuration

DGX agent

arXiv:2604.09678v1 Announce Type: cross Abstract: As agentic network management gains popularity, there is a critical need for evaluation frameworks that transcend static, one-shot testing. To address

model-releasesarxiv-cs-ai
14 Apr 2026
Hardware

Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse

DGX agent

arXiv:2511.00413v4 Announce Type: replace Abstract: Agentic large language model (LLM) training often involves multi-turn interaction trajectories that branch into multiple execution paths due to conc

hardwarearxiv-cs-lg
14 Apr 2026
Agents

Working Paper: Towards Schema-based Learning from a Category-Theoretic Perspective

DGX agent

arXiv:2604.10589v1 Announce Type: new Abstract: We introduce a hierarchical categorical framework for Schema-Based Learning (SBL) structured across four interconnected levels. At the schema level, a f

agentsarxiv-cs-ai
14 Apr 2026
Agents

ActionNex: A Virtual Outage Manager for Cloud Computing

DGX agent

arXiv:2604.03512v2 Announce Type: replace Abstract: Outage management in large-scale cloud operations remains heavily manual, requiring rapid triage, cross-team coordination, and experience-driven dec

agentsarxiv-cs-ai
13 Apr 2026
Model Releases

Adaptive Tuning of Parameterized Traffic Controllers via Multi-Agent Reinforcement Learning

DGX agent

arXiv:2512.07417v2 Announce Type: replace Abstract: Effective traffic control is essential for mitigating congestion in transportation networks. Conventional traffic management strategies, including r

model-releasesarxiv-cs-lg
13 Apr 2026
Safety

AgentSociety: Large-Scale Simulation of LLM-Driven Generative Agents Advances Understanding of Human Behaviors and Society

DGX agent

arXiv:2502.08691v2 Announce Type: replace-cross Abstract: Understanding human behavior and society is a central focus in social sciences, with the rise of generative social science marking a significa

safetyarxiv-cs-ai
13 Apr 2026
Safety

Semantic Intent Fragmentation: A Single-Shot Compositional Attack on Multi-Agent AI Pipelines

DGX agent

arXiv:2604.08608v1 Announce Type: cross Abstract: We introduce Semantic Intent Fragmentation (SIF), an attack class against LLM orchestration systems where a single, legitimately phrased request cause

safetyarxiv-cs-ai
13 Apr 2026
Agents

Strategic Algorithmic Monoculture:Experimental Evidence from Coordination Games

DGX agent

arXiv:2604.09502v1 Announce Type: new Abstract: AI agents increasingly operate in multi-agent environments where outcomes depend on coordination. We distinguish primary algorithmic monoculture -- base

agentsarxiv-cs-ai
13 Apr 2026
Model Releases

Structured Uncertainty guided Clarification for LLM Agents

DGX agent

arXiv:2511.08798v2 Announce Type: replace-cross Abstract: LLM agents with tool-calling capabilities often fail when user instructions are ambiguous or incomplete, leading to incorrect invocations and

model-releasesarxiv-cs-ai
13 Apr 2026
Safety

VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning

DGX agent

arXiv:2604.09508v1 Announce Type: cross Abstract: Visual Retrieval-Augmented Generation (VRAG) empowers Vision-Language Models to retrieve and reason over visually rich documents. To tackle complex qu

safetyarxiv-cs-ai
13 Apr 2026
Safety

An Agentic Evaluation Architecture for Historical Bias Detection in Educational Textbooks

DGX agent

arXiv:2604.07883v1 Announce Type: cross Abstract: History textbooks often contain implicit biases, nationalist framing, and selective omissions that are difficult to audit at scale. We propose an agen

safetyarxiv-cs-cl
10 Apr 2026
Model Releases

Don't Overthink It: Inter-Rollout Action Agreement as a Free Adaptive-Compute Signal for LLM Agents

DGX agent

arXiv:2604.08369v1 Announce Type: cross Abstract: Inference-time compute scaling has emerged as a powerful technique for improving the reliability of large language model (LLM) agents, but existing me

model-releasesarxiv-cs-cl
10 Apr 2026
Safety

Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization

DGX agent

arXiv:2604.07165v1 Announce Type: new Abstract: Reinforcement learning for Large Language Model agents is often hindered by sparse rewards in multi-step reasoning tasks. Existing approaches like Group

safetyarxiv-cs-ai
10 Apr 2026
Model Releases

SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills

DGX agent

arXiv:2604.06550v1 Announce Type: cross Abstract: OpenClaw's ClawHub marketplace hosts over 13,000 community-contributed agent skills, and between 13% and 26% of them contain security vulnerabilities

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

Strategic Persuasion with Trait-Conditioned Multi-Agent Systems for Iterative Legal Argumentation

DGX agent

arXiv:2604.07028v1 Announce Type: cross Abstract: Strategic interaction in adversarial domains such as law, diplomacy, and negotiation is mediated by language, yet most game-theoretic models abstract

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-Auditing

DGX agent

arXiv:2604.08401v1 Announce Type: cross Abstract: In large language model (LLM) agents, reasoning trajectories are treated as reliable internal beliefs for guiding actions and updating memory. However

model-releasesarxiv-cs-cl
10 Apr 2026
Model Releases

VisCoder2: Building Multi-Language Visualization Coding Agents

DGX agent

arXiv:2510.23642v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently enabled coding agents capable of generating, executing, and revising visualization code. However, e

model-releasesarxiv-cs-ai
10 Apr 2026
Safety

WebExpert: domain-aware web agents with critic-guided expert experience for high-precision search

DGX agent

arXiv:2604.06177v1 Announce Type: cross Abstract: Specialized web tasks in finance, biomedicine, and pharmaceuticals remain challenging due to missing domain priors: queries drift, evidence is noisy,

safetyarxiv-cs-ai
10 Apr 2026
Model Releases

Do LLMs Beat Nash? Testing Decentralized Coordination in Self-Play Multi-Agent Games

DGX agent

arXiv:2608.12547v1 Announce Type: cross Abstract: Large language model agents deployed without a central controller are often assumed to require communication to coordinate their actions. We ask what

model-releasesarxiv-cs-ro
14 Aug 2026
Model Releases

Doctorina MedBench: A Dialogue-Based Benchmark and Evaluation Framework for Agent-Based Medical AI

DGX agent

arXiv:2603.25821v3 Announce Type: replace-cross Abstract: We present Doctorina MedBench, an evaluation framework for agent-based medical AI based on the simulation of physician-patient interactions. U

model-releasesarxiv-cs-ai
14 Aug 2026
Safety

Intern-S2-Preview: Scientific Agentic Foundation Model

DGX agent

arXiv:2608.13505v1 Announce Type: cross Abstract: Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific t

safetyarxiv-cs-cl
14 Aug 2026
Model Releases

PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives

DGX agent

arXiv:2608.13552v1 Announce Type: new Abstract: Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consis

model-releasesarxiv-cs-cv
14 Aug 2026
Safety

Reconcile Once, Write Anytime: A Trust-Tiered Librarian and a Multi-Agent Writer for Drift-Free, Point-in-Time Research

DGX agent

arXiv:2608.12984v1 Announce Type: cross Abstract: Long-form research reports generated by large language models drift, contradict themselves, and lose provenance: the same metric appears with differen

safetyarxiv-cs-cl
14 Aug 2026
Research

SynWeaver: Website-Prior Task and Trajectory Co-Synthesis for Web Agents

DGX agent

arXiv:2608.12429v1 Announce Type: cross Abstract: Web agents often struggle to generalize to unseen websites because they lack website-specific supervision. Recent exploration-based data synthesis met

researcharxiv-cs-ai
14 Aug 2026
Model Releases

Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs

DGX agent

arXiv:2608.11232v1 Announce Type: cross Abstract: Evaluating LLM coding agents in algorithmic trading is difficult because static benchmarks risk data contamination and numerical backtest outputs requ

model-releasesarxiv-cs-ai
13 Aug 2026
Agents

Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents

DGX agent

arXiv:2608.11552v1 Announce Type: cross Abstract: Uncertainty quantification (UQ) methods for language models are typically evaluated on single-turn outputs, where uncertainty is attached to one gener

agentsarxiv-cs-ai
13 Aug 2026
← Previous
1…7374757677…236
Next →