AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,153 results
Model Releases

EMAS: Stabilizing Multi-Agent System Evolution through Evidence-Guided Revision

DGX agent

arXiv:2608.07196v1 Announce Type: new Abstract: Many methods for automated multi-agent system design optimize prompts and topologies during an initial design stage and then deploy the resulting system

model-releasesarxiv-cs-ai
10 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Local Ai

Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay

DGX agent

arXiv:2608.05784v1 Announce Type: new Abstract: Computer-use agents pay full frontier inference to re-derive routines their user has already performed, because an agent's memory today records what the

local-aiarxiv-cs-ai
7 Aug 2026
Agents

Agentic Software Issue Resolution with Large Language Models: A Survey

DGX agent

arXiv:2512.22256v2 Announce Type: replace-cross Abstract: Software issue resolution aims to address real-world issues in software repositories based on natural language descriptions provided by users,

agentsarxiv-cs-ai
7 Aug 2026
Agents

ASTELD: A Six-Axis Classification Framework for Autonomous AI Agents - Design, Evaluation, and an OpenClaw Case Study

DGX agent

arXiv:2608.05201v1 Announce Type: cross Abstract: Autonomous AI agent platforms differ substantially in architecture, security, tool integration, execution, autonomy, and deployment, yet the field lac

agentsarxiv-cs-ai
7 Aug 2026
Agents

BioAgent Bench: An AI Agent Evaluation Suite for Bioinformatics

DGX agent

arXiv:2601.21800v4 Announce Type: replace Abstract: We introduce BioAgent Bench, an evaluation suite designed for measuring the performance and robustness of AI agents in common bioinformatics tasks.

agentsarxiv-cs-ai
7 Aug 2026
Agents

Learning Globally Reusable Skills for Coding Agents

DGX agent

arXiv:2608.06153v1 Announce Type: cross Abstract: Automated skill evolution enables Large Language Model (LLM) agents to continuously improve without expensive retraining. However, existing approaches

agentsarxiv-cs-ai
7 Aug 2026
Agents

AgentAntibody: An Adaptive Immune System for Defending LLM Agents against Prompt Injection

DGX agent

arXiv:2608.04053v1 Announce Type: cross Abstract: Prompt injection remains a critical threat to LLM agents, yet existing defenses treat each task as a self-contained problem, independent of previous e

agentsarxiv-cs-ai
6 Aug 2026
Agents

Interoceptive Attention as Dynamic Homeostatic Prioritization in a Foraging Agent

DGX agent

arXiv:2608.04232v1 Announce Type: new Abstract: Biological systems must regulate competing needs under limited perceptual bandwidth, where sharpening one estimate costs the capacity to sharpen the oth

agentsarxiv-cs-ai
6 Aug 2026
Agents

ReCodeAgent: A Multi-agent Workflow for Language-Agnostic Translation and Validation of Large-Scale Repositories

DGX agent

arXiv:2604.07341v2 Announce Type: replace-cross Abstract: Most repository-level code translation and validation techniques have been evaluated on a single source-target programming language (PL) pair,

agentsarxiv-cs-lg
6 Aug 2026
Agents

An Actionable Diagnosis of Multilingual, Multi-Agent Planning Failures

DGX agent

arXiv:2608.03735v1 Announce Type: cross Abstract: Multilingual multi-agent systems exhibit substantial degradation beyond English, yet prior work rarely identifies how task-critical information is los

agentsarxiv-cs-cl
5 Aug 2026
Agents

Hybrid LLM-Augmented Reinforcement Learning Agents for Complex Sequential Decision Tasks

DGX agent

arXiv:2608.03502v1 Announce Type: new Abstract: Large Language Models (LLMs) have recently shown strong capabilities in reasoning, planning, and tool-use, enabling new forms of autonomous agents. Howe

agentsarxiv-cs-ai
5 Aug 2026
Model Releases

S^3: Improving Agent Safety through Multi-Stage Defense

DGX agent

arXiv:2608.02683v1 Announce Type: cross Abstract: Large Language Model (LLM) agents rely on multi-stage agentic workflows, with stages such as memory, planning, and tool execution, to accomplish compl

model-releasesarxiv-cs-ai
5 Aug 2026
Agents

Steganalysis of Adaptive Covert Collusion in Tool-Using Agent Populations: A Black-Box, Cross-Principal Approach

DGX agent

arXiv:2608.02698v1 Announce Type: cross Abstract: Tool-using agents built on large language models (LLMs) are increasingly deployed not by a single operator but by many, side by side on shared infrast

agentsarxiv-cs-ai
5 Aug 2026
Agents

Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance

DGX agent

arXiv:2608.03403v1 Announce Type: new Abstract: The performance bottleneck of agents is increasingly shifting from model capability to the robustness of their execution processes. Tools play a central

agentsarxiv-cs-ai
5 Aug 2026
Model Releases

UrbanAgent: A Tool-Augmented Agent for Cross-System Urban Tasks

DGX agent

arXiv:2608.03018v1 Announce Type: new Abstract: Modern cities rely on an increasing number of digital services to operate, but residents' daily needs are still difficult to meet. Services are fragment

model-releasesarxiv-cs-ai
5 Aug 2026
Agents

Agentic Bayesian Optimization through Surrogate-Augmented Autoresearch

DGX agent

arXiv:2608.00316v1 Announce Type: new Abstract: Bayesian optimization (BO) has become the standard tool for sample-efficient optimization and owes its efficiency to uncertainty-aware search driven by

agentsarxiv-cs-lg
4 Aug 2026
Model Releases

HarnessCompass: Guiding Automatic Harness Evolution toward Generalizable and Effective Agent Harnesses

DGX agent

arXiv:2608.01918v1 Announce Type: cross Abstract: Harness design plays a critical role in agent performance by shaping how large language models (LLMs) perceive, reason over, and act within executable

model-releasesarxiv-cs-cl
4 Aug 2026
Agents

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving

DGX agent

arXiv:2512.10739v3 Announce Type: replace Abstract: Large Reasoning Models (LRMs) have expanded the mathematical reasoning frontier through Chain-of-Thought (CoT) techniques and Reinforcement Learning

agentsarxiv-cs-cl
4 Aug 2026
Model Releases

Prompt-Induced Waste in Large Reasoning Models: A Preregistered Two-Harness Benchmark of Coding Agents

DGX agent

arXiv:2608.01347v1 Announce Type: new Abstract: Large reasoning models used as coding agents incur costs from deliberation, tool calls, and repeated agent turns, yet the causal effect of prompt wordin

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers

DGX agent

arXiv:2607.29626v1 Announce Type: new Abstract: As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly i

model-releasesarxiv-cs-ai
3 Aug 2026
Model Releases

CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents

DGX agent

arXiv:2607.19338v2 Announce Type: replace Abstract: Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an incorrect an

model-releasesarxiv-cs-ai
3 Aug 2026
Safety

MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action Distillation

DGX agent

arXiv:2607.29320v1 Announce Type: new Abstract: Graphical user interface (GUI) agents based on large language models are increasingly deployed across mobile, web, and desktop environments. However, ex

safetyarxiv-cs-ai
3 Aug 2026
Hardware

AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolution

DGX agent

arXiv:2607.26661v1 Announce Type: new Abstract: Ascend C operator optimization is critical for NPU (Neural Processing Unit) inference performance but requires deep hardware expertise.While large langu

hardwarearxiv-cs-ai
31 Jul 2026
Model Releases

ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents

DGX agent

arXiv:2607.28037v1 Announce Type: new Abstract: As LLM-based agents are deployed in complex, multi-step workflows, a critical evaluation gap has emerged: most existing benchmarks judge only final outc

model-releasesarxiv-cs-lg
31 Jul 2026
Model Releases

Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?

DGX agent

arXiv:2607.26367v1 Announce Type: new Abstract: An important skill in theoretical physics is to recognize when a new problem can be transformed into a known model. We study this skill as an AI-agent t

model-releasesarxiv-cs-ai
31 Jul 2026
Safety

Inference-Time Agentic Decision Rules Beat Longer Evolving Search for Multi-Image Medical Reasoning

DGX agent

arXiv:2607.27564v1 Announce Type: new Abstract: Multi-image medical VQA is not merely a prompt-length problem; it is a fundamental challenge of agentic decision-making. Medical vision-language agents

safetyarxiv-cs-cv
31 Jul 2026
Safety

LabEvolver: Training-Free Experience Evolution for Safe and Grounded Wet-Lab Agents

DGX agent

arXiv:2607.27690v1 Announce Type: new Abstract: We introduce LabEvolver, a training-free framework that equips safe and grounded wet-lab agents with episodic memory from execution experience. LabEvolv

safetyarxiv-cs-ro
31 Jul 2026
Agents

MemTxn: A Transaction Boundary for Source-Supported Updates and Complete-State Recovery in Agent Memory

DGX agent

arXiv:2607.27834v1 Announce Type: cross Abstract: Persistent memory lets long-running large language model agents reuse information across sessions and tasks. Yet errors in writable memory can persist

agentsarxiv-cs-cl
31 Jul 2026
Local Ai

AgentGUI: An Interface for Observing and Steering Long-Running AI Agents

DGX agent

arXiv:2607.26300v1 Announce Type: new Abstract: AI agents are increasingly adept at tackling complex, long-running tasks. With the rapid surge of autonomous capabilities, human oversight is systematic

local-aiarxiv-cs-cl
30 Jul 2026
Safety

AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents

DGX agent

arXiv:2607.26998v1 Announce Type: cross Abstract: Large language model (LLM) agents automate penetration testing through an observation-action loop, selecting actions based on observations returned by

safetyarxiv-cs-cl
30 Jul 2026
Agents

Agent Skills Matter: Inferring Proprietary Skills from Execution Trajectories

DGX agent

arXiv:2607.25560v1 Announce Type: new Abstract: Agent skills package reusable procedures that improve downstream performance. Their lightweight, portable form enables marketplace monetization and priv

agentsarxiv-cs-ai
29 Jul 2026
Safety

AlphaCrafter: Harnessing Multi-Agent Workflows for Cross-Sectional Quantitative Trading

DGX agent

arXiv:2605.05580v2 Announce Type: replace Abstract: Quantitative trading agents have demonstrated substantial promise in automating factor discovery, signal aggregation, and portfolio execution. Howev

safetyarxiv-cs-ai
29 Jul 2026
Agents

Distributing Security Controls Through Harness Engineering

DGX agent

arXiv:2607.25890v1 Announce Type: new Abstract: AI coding agents are being adopted at historic speed, yet security and risk concerns remain the primary barrier to scaling agentic AI across organizatio

agentsarxiv-cs-ai
29 Jul 2026
Agents

Runtime Uncertainty Monitoring for LLM-Based Multi-Agent Systems Using Bayesian Networks

DGX agent

arXiv:2607.25877v1 Announce Type: new Abstract: This paper investigates how multi-agent systems (MAS)-based on large language models (LLMs) can support actuarial risk modelling, with a particular focu

agentsarxiv-cs-ai
29 Jul 2026
Agents

Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction

DGX agent

arXiv:2607.25718v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly rely on invoking external tools to complete real-world tasks. Tool retrieval, which selects a small tas

agentsarxiv-cs-ai
29 Jul 2026
Agents

VLD-RAG: Agentic Vision-Language Retrieval-Augmented Generation for Long, Visually-Rich Multi-Page Documents

DGX agent

arXiv:2607.24748v1 Announce Type: cross Abstract: Visually-rich documents such as reports, slides, and manuals often distribute the evidence needed to answer a question across multiple pages, mixing t

agentsarxiv-cs-ai
29 Jul 2026
Local Ai

Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents

DGX agent

arXiv:2607.22688v1 Announce Type: new Abstract: Post-training agents for automated AI research requires optimizing not only model parameters, but also the runtime harness that shapes how research traj

local-aiarxiv-cs-ai
28 Jul 2026
Model Releases

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios

DGX agent

arXiv:2607.23722v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as agents that interact with stateful environments over multiple steps: gathering hidden informat

model-releasesarxiv-cs-ai
28 Jul 2026
Safety

MemChain: Learning Interpretable Memory Traces for Memory-Augmented LLM Agents

DGX agent

arXiv:2607.24097v1 Announce Type: new Abstract: Memory-augmented LLM agents typically answer queries by retrieving relevant memories and feeding them directly to an answer model. This retrieval-as-evi

safetyarxiv-cs-ai
28 Jul 2026
Model Releases

SCAIR: Schema-Conditioned Agentic Iterative Reasoning for Enterprise Knowledge Graphs

DGX agent

arXiv:2607.22571v1 Announce Type: new Abstract: Knowledge Graph-based Retrieval-Augmented Generation (KG-RAG) enables natural language interaction with structured enterprise knowledge, yet existing ag

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

SIREN: Towards End-to-End Extreme-Weather Early Warning with Experience-Grounded LLM Agents

DGX agent

arXiv:2607.24588v1 Announce Type: new Abstract: Early warning of extreme weather is essential for mitigating the societal, economic, and environmental risks posed by hazardous weather events. However,

model-releasesarxiv-cs-ai
28 Jul 2026
Agents

Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation Detection

DGX agent

arXiv:2607.18080v2 Announce Type: replace-cross Abstract: Multimodal video misinformation detection is commonly formulated as a holistic video-understanding task, where the entire video and its associ

agentsarxiv-cs-ai
28 Jul 2026
Safety

Spectral Dynamics of Semantic Drift in Clinical Multi-Agent Language Model Networks

DGX agent

arXiv:2607.22758v1 Announce Type: cross Abstract: The integration of iterative LLMs within multi-agent diagnostic frameworks requires a rigorous quantitative reevaluation of underlying communication t

safetyarxiv-cs-ai
28 Jul 2026
Model Releases

DBA-Bench: A Production-Fidelity Benchmark for LLM-Based Database Operations Agents

DGX agent

arXiv:2607.22165v1 Announce Type: cross Abstract: LLM-based database agents show promise, but differing task scopes, testbeds, and metrics hinder comparison. We identify four gaps between evaluation a

model-releasesarxiv-cs-cl
27 Jul 2026
Agents

Towards Reducing Foreign Language Anxiety Using Level-Appropriate Embodied Conversational Agents

DGX agent

arXiv:2607.21887v1 Announce Type: cross Abstract: Foreign language anxiety (FLA) can be a major barrier to second language acquisition (SLA), especially in conversational contexts. With the proliferat

agentsarxiv-cs-cl
27 Jul 2026
Model Releases

InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents

DGX agent

arXiv:2607.20468v1 Announce Type: new Abstract: AI agents are increasingly used to automate research and development tasks, yet existing benchmarks typically evaluate them on prescribed workflows or n

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

OpenForgeRL: Train Harness-native Agents in Any Environment

DGX agent

arXiv:2607.21557v1 Announce Type: new Abstract: Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to e

model-releasesarxiv-cs-ai
24 Jul 2026
Agents

A Framework of User Experience Principles for Human-AI Agent Interaction in the Workplace

DGX agent

arXiv:2607.19941v1 Announce Type: cross Abstract: As AI agents become integral to business workflows, establishing guiding user experience (UX) principles is crucial for ensuring user trust and succes

agentsarxiv-cs-ai
23 Jul 2026
← Previous
1…3031323334…233
Next →