AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,040 results
Agents

Automating and Scaling Behavioral Scientific Research on AI Agents

DGX agent

arXiv:2608.10030v1 Announce Type: new Abstract: As AI agents are increasingly deployed in complex environments, understanding their behaviors becomes critical. Yet behavioral scientific research on AI

agentsarxiv-cs-ai
12 Aug 2026
Agents
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Do Personalized Skills Help Coding Agents? An Empirical Study of Developer Interaction Histories

DGX agent

arXiv:2608.10319v1 Announce Type: cross Abstract: Large language model (LLM)-powered agents have rapidly evolved from code-completion tools into solvers of complex software engineering tasks. As devel

agentsarxiv-cs-ai
12 Aug 2026
Agents

FlowScout: From Execution Feedback to Reliable Tool-Using Agent Workflows

DGX agent

arXiv:2608.10039v1 Announce Type: new Abstract: Agentic workflows have become an important abstraction for building reliable LLM-based automation systems by organizing large language models (LLMs), to

agentsarxiv-cs-lg
12 Aug 2026
Agents

SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent

DGX agent

arXiv:2608.07449v1 Announce Type: new Abstract: LLM agents increasingly adapt to recurring tasks by accumulating procedural knowledge in skills. These skills are lightweight, reusable textual artifact

agentsarxiv-cs-ai
10 Aug 2026
Safety

Search-Aided Joint Agent-Environment Reinforcement Learning for Robust Lifelong Multi-Agent Path Finding with Rotations

DGX agent

arXiv:2608.05588v1 Announce Type: cross Abstract: Lifelong Multi-Agent Path Finding (LMAPF) requires repeatedly planning collision-free paths for agents that continuously receive new goals upon reachi

safetyarxiv-cs-ai
7 Aug 2026
Model Releases

A Survey of Agent Memory in the Second Half: Towards Self-Evolving and Long-Horizon Agents

DGX agent

arXiv:2602.06052v4 Announce Type: replace-cross Abstract: Research in artificial intelligence is shifting from model innovations and benchmark scores towards problem definition and rigorous real-world

model-releasesarxiv-cs-ai
6 Aug 2026
Model Releases

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation

DGX agent

arXiv:2608.03166v1 Announce Type: new Abstract: Role-Playing Language Agents (RPLAs) are increasingly deployed in high-stakes applications such as healthcare assistance, customer support, and educatio

model-releasesarxiv-cs-ai
5 Aug 2026
Agents

Autoreflection: How Agentic Strange Loops Turn Human Culture into AI Infrastructure

DGX agent

arXiv:2608.03800v1 Announce Type: cross Abstract: An LLM-based agent is a loop that reads itself. Agentic frameworks externalize identity, memory, and disposition into editable files. The agent loads

agentsarxiv-cs-ai
5 Aug 2026
Agents

Practical Online KV Cache Compaction for LLM Agents: An Empirical Study

DGX agent

arXiv:2608.00902v1 Announce Type: new Abstract: LLM agents accumulate long trajectories of reasoning steps, tool calls, and environment feedback, making the KV cache a major inference bottleneck. KV c

agentsarxiv-cs-cl
4 Aug 2026
Model Releases

UrbanDS: A Graph-Guided LLM Multi-Agent System for Data-Intensive Urban Tasks

DGX agent

arXiv:2607.26724v1 Announce Type: new Abstract: Large language model (LLM) agents have been widely applied in automating data science tasks. However, existing methods typically rely on a limited set o

model-releasesarxiv-cs-ai
31 Jul 2026
Safety

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops

DGX agent

arXiv:2607.25152v1 Announce Type: new Abstract: Long-running autonomous agents plan, act, and judge their own completion without human intervention. When an agent grades its own work, self-evaluation

safetyarxiv-cs-ai
29 Jul 2026
Agents

A Vocabulary for Multi-Agent Automated Research Systems

DGX agent

arXiv:2607.22682v1 Announce Type: new Abstract: We introduce a vocabulary for automated research systems built from one or more agents to make their design choices easier to describe and compare. The

agentsarxiv-cs-ai
28 Jul 2026
Agents

Agentic Reward Modeling: Verifying GUI Agent via Progressive Trajectory-Grounded Interaction

DGX agent

arXiv:2602.00575v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) provides a promising pathway for continuously advancing GUI agents, yet existing reward modeli

agentsarxiv-cs-ro
28 Jul 2026
Agents

Plans Work in Mysterious Ways: Evaluating a Plan Mode for Spreadsheet Agents

DGX agent

arXiv:2607.23670v1 Announce Type: cross Abstract: Plan Modes have become standard features in agentic programming tools, allowing users to gain transparency and control by working with the agent to de

agentsarxiv-cs-ai
28 Jul 2026
Agents

The Best Programming Language for Tokenmaxxing: An Investigation of Coding Agent Behavior Across Programming Languages

DGX agent

arXiv:2607.22807v1 Announce Type: cross Abstract: Although coding agents are now very effective in a variety of programming languages, this paper first shows that the cost (in tokens) can very signifi

agentsarxiv-cs-cl
28 Jul 2026
Model Releases

The Scaffold Effect in Coding Agents: Harness Choice as a Hidden Variable in Coding-Agent Evaluation

DGX agent

arXiv:2607.22585v1 Announce Type: new Abstract: Public leaderboards for coding agents typically rank systems by model name and pass rate, while the surrounding harness (the scaffold that issues tools,

model-releasesarxiv-cs-ai
28 Jul 2026
Safety

Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Mode

DGX agent

arXiv:2607.22083v1 Announce Type: cross Abstract: We present Nanbeige4.2-3B, a compact general agentic model with 3B non-embedding parameters. It delivers strong performance across code-agent, office-

safetyarxiv-cs-cl
27 Jul 2026
Safety

Learning Red Agent Policy from Observations for Neurosymbolic Autonomous Cyber Agents

DGX agent

arXiv:2606.18223v2 Announce Type: replace-cross Abstract: With sophisticated cyber-attacks becoming increasingly prevalent, modern networks require intelligent autonomous cyber-defense agents trained

safetyarxiv-cs-ai
16 Jul 2026
Agents

Fin-Analyst at FinMMEval 2026 Task 3: A Live Hybrid Trading Agent with LLM Specialists and Rule-Based Signals

DGX agent

arXiv:2607.12233v1 Announce Type: cross Abstract: Large language model (LLM) trading agents show promising performance in equity markets, yet remain narrowly focused on US equities with little evidenc

agentsarxiv-cs-ai
15 Jul 2026
Model Releases

Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents

DGX agent

arXiv:2607.08716v1 Announce Type: new Abstract: In long-horizon tasks, decision-relevant state is often scattered across an expanding trajectory, while the action agent must surface it and act. As tra

model-releasesarxiv-cs-ai
10 Jul 2026
Local Ai

When Agents Go Rogue: Activation-Based Detection of Malicious Behaviors in Multi-Agent Systems

DGX agent

arXiv:2607.06807v1 Announce Type: cross Abstract: While enabling effective collaboration on complex tasks, LLM-based Multi-Agent Systems (MAS) face critical security challenges due to vulnerabilities

local-aiarxiv-cs-ai
9 Jul 2026
Agents

MechMath Agent Team: LLM Driven Agents for Mathematical Research

DGX agent

arXiv:2607.04394v1 Announce Type: new Abstract: AI reasoning has become a central focus in contemporary artificial intelligence, largely driven by the success of large language models. However, mathem

agentsarxiv-cs-ai
7 Jul 2026
Model Releases

PolicyGuard: A Dialogue-Grounded Sub-Agent Verifier for Policy Adherence in LLM Agents

DGX agent

arXiv:2606.29225v1 Announce Type: new Abstract: LLM agents handle user requests on behalf of organizations through tool calls and must follow the company policies stated in their system prompts. Prior

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent

DGX agent

arXiv:2606.30616v1 Announce Type: new Abstract: We introduce Agents-A1, a 35B Mixture-of-Experts Agentic Model that reaches trillion-parameter-level performance by scaling the agent horizon. We invest

model-releasesarxiv-cs-cl
30 Jun 2026
Hardware

SwarmX: Agentic Scheduling for Low-Latency Agentic Systems

DGX agent

arXiv:2606.21401v2 Announce Type: replace-cross Abstract: Agentic AI applications compose multiple model calls and tool executions, creating new scheduling challenges for GPU-CPU clusters. Their infer

hardwarearxiv-cs-ai
30 Jun 2026
Model Releases

Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models

DGX agent

arXiv:2606.18142v3 Announce Type: replace Abstract: AI agents are moving from advisors to actors, booking travel, planning menus, and running procurement on behalf of users. Existing benchmarks for AI

model-releasesarxiv-cs-ai
29 Jun 2026
Safety

BARD-MARL: Byzantine-Agent Detection for Learned Communication in Multi-Agent Reinforcement Learning

DGX agent

arXiv:2606.20701v1 Announce Type: cross Abstract: Learned communication improves coordination in cooperative multi-agent reinforcement learning, but it also creates a trust problem: a trained policy m

safetyarxiv-cs-lg
23 Jun 2026
Model Releases

Divide and Cooperate: Role-Decomposed Multi-Agent LLM Training with Cross-Agent Learning Signals

DGX agent

arXiv:2606.10684v1 Announce Type: cross Abstract: Modern language agents which perform multi-step reasoning have shown strong performance in knowledge-intensive question answering. However, existing a

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents

DGX agent

arXiv:2512.23128v2 Announce Type: replace-cross Abstract: Web-based agents powered by large language models are increasingly used for tasks such as email management or professional networking. Their r

model-releasesarxiv-cs-ai
8 Jun 2026
Agents

RAINO: Anchoring Agents in Reality, A Systematic Review and Conceptual Framework for Realism in Agent-Based Modelling

DGX agent

arXiv:2606.05167v1 Announce Type: cross Abstract: Realism is a central yet seemingly under-theorized concept in Agent-Based Modelling. This paper presents a Systematic Literature Review, aiming to ide

agentsarxiv-cs-ai
6 Jun 2026
Model Releases

Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents

DGX agent

arXiv:2606.04874v1 Announce Type: new Abstract: Planning is central to LLM agents: before acting, an agent must decompose goals, select tools, reason over constraints, and decide when a task is infeas

model-releasesarxiv-cs-cl
4 Jun 2026
Safety

From Agent Traces to Trust: Evidence Tracing and Execution Provenance in LLM Agents

DGX agent

arXiv:2606.04990v1 Announce Type: cross Abstract: Large language model (LLM)-based agents increasingly solve complex tasks by interacting with external tools, retrieval systems, memory modules, enviro

safetyarxiv-cs-ai
4 Jun 2026
Model Releases

Proof-Carrying Agent Actions: Model-Agnostic Runtime Governance for Heterogeneous Agent Systems

DGX agent

arXiv:2606.04104v1 Announce Type: cross Abstract: Agent systems execute through runtimes with very different control points: local coding tools, framework SDKs, managed agent platforms, API gateways,

model-releasesarxiv-cs-ai
4 Jun 2026
Agents

Agentic-J: An AI Agent for Biological Microscopy Image Analysis

DGX agent

arXiv:2606.02080v1 Announce Type: cross Abstract: Biological image analysis increasingly demands integration across heterogeneous tools, programming environments, and domain knowledge that few researc

agentsarxiv-cs-ai
2 Jun 2026
Agents

BAGEN: Are LLM Agents Budget-Aware?

DGX agent

arXiv:2606.00198v1 Announce Type: cross Abstract: While agents are increasingly spending more resources, today agent cost is mostly measured only after execution. A Budget-Aware Agent (BAGEN) should t

agentsarxiv-cs-ai
2 Jun 2026
Safety

COMAP: Co-Evolving World Models and Agent Policies for LLM Agents

DGX agent

arXiv:2606.02372v1 Announce Type: new Abstract: Equipping language agents with world models enables them to anticipate environment dynamics and evaluate candidate actions before execution. However, ex

safetyarxiv-cs-ai
2 Jun 2026
Model Releases

Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories

DGX agent

arXiv:2606.02060v1 Announce Type: new Abstract: Deep-research agents solve tasks through long trajectories of search, tool use, evidence inspection, and answer synthesis. Evaluation based on final ans

model-releasesarxiv-cs-ai
2 Jun 2026
Agents

Catalyst-Agent: Autonomous heterogeneous catalyst screening with an LLM Agent

DGX agent

arXiv:2603.01311v2 Announce Type: replace Abstract: The discovery of novel catalysts tailored for particular applications is a major challenge for the twenty-first century. Traditional methods for thi

agentsarxiv-cs-cl
29 May 2026
Hardware

Long Live the Librarian! A Persistent Search Sub-Agent for Energy-Efficient Multi-Agent Software Engineering Systems

DGX agent

arXiv:2605.27787v1 Announce Type: cross Abstract: Multi-agent systems (MAS) have substantially advanced autonomous software engineering (SWE), but their growing inference energy demands raise sustaina

hardwarearxiv-cs-cl
28 May 2026
Agents

Agent-as-Peer-Debriefer: A Multi-Agent Framework with Perspective-Based Refinement for Qualitative Analysis

DGX agent

arXiv:2605.24600v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for qualitative data analysis (QDA), yet their outputs often miss the depth and nuance of human analy

agentsarxiv-cs-ai
26 May 2026
Model Releases

Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning

DGX agent

arXiv:2602.10090v3 Announce Type: replace Abstract: Recent advances in large language model (LLM) have empowered autonomous agents to perform multi-turn interactions with tools and environments. Howev

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks

DGX agent

arXiv:2505.24876v2 Announce Type: replace-cross Abstract: Deep reasoning is fundamental for solving complex tasks, especially in vision-centric scenarios that demand sequential, multimodal understandi

model-releasesarxiv-cs-cl
26 May 2026
Agents

MARGIN: Runtime Confidence Calibration for Multi-Agent Foundation Model Coordination

DGX agent

arXiv:2605.22949v1 Announce Type: new Abstract: Foundation model agents increasingly operate in multi-agent deployments where a coordinator must decide which agent's response to trust. The standard ap

agentsarxiv-cs-lg
25 May 2026
Agents

Agent JIT Compilation for Latency-Optimizing Web Agent Planning and Scheduling

DGX agent

arXiv:2605.21470v1 Announce Type: new Abstract: Computer-use agents (CUA) automate tasks specified with natural language such as 'order the cheapest item from Taco Bell' by generating sequences of cal

agentsarxiv-cs-lg
21 May 2026
Agents

Agentic Trading: When LLM Agents Meet Financial Markets

DGX agent

arXiv:2605.19337v1 Announce Type: new Abstract: A growing body of work explores how Large Language Models (LLMs) can be embedded in trading systems as agents that perceive market information, retrieve

agentsarxiv-cs-ai
20 May 2026
Agents

IR-Agent: Expert-Inspired LLM Agents for Structure Elucidation from Infrared Spectra

DGX agent

arXiv:2508.16112v2 Announce Type: replace Abstract: Spectral analysis provides crucial clues for the elucidation of unknown materials. Among various techniques, infrared spectroscopy (IR) plays an imp

agentsarxiv-cs-ai
20 May 2026
Agents

Agents for Experiments, Experiments for Agents: A Design Grammar for AI-Enabled Experimental Science

DGX agent

arXiv:2605.17746v1 Announce Type: new Abstract: AI systems are becoming active participants in organizational and knowledge work. They increasingly interact with humans, coordinate workflows, and oper

agentsarxiv-cs-ai
19 May 2026
Agents

Orchard: An Open-Source Agentic Modeling Framework

DGX agent

arXiv:2605.15040v1 Announce Type: new Abstract: Agentic modeling aims to transform LLMs into autonomous agents capable of solving complex tasks through planning, reasoning, tool use, and multi-turn in

agentsarxiv-cs-ai
15 May 2026
← Previous
1…678910…230
Next →