AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,153 results
Safety

Similarity Gates Approve Reversals: A Validity Audit of Embedding-Cosine Thresholds in Agent Systems

DGX agent

arXiv:2608.10216v1 Announce Type: cross Abstract: Agent frameworks ship quality gates that compare text blocks by embedding-cosine similarity and decide at a fixed cutoff. Deduplication filters, seman

safetyarxiv-cs-ai
12 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

A Unified Issue Resolution Benchmark for Requirement Clarification, Planning, and Code Generation for Coding Agents

DGX agent

arXiv:2608.09072v1 Announce Type: cross Abstract: Large language model-powered coding agents are increasingly used to modify existing code repositories, for example, by adding features or fixing bugs.

model-releasesarxiv-cs-ai
11 Aug 2026
Agents

Agentic Auto-Research is Fuzz Testing

DGX agent

arXiv:2608.09855v1 Announce Type: new Abstract: Autonomous research agents can generate experiments faster than researchers can validate them. Researchers have responded by scaling the proposer and ra

agentsarxiv-cs-ai
11 Aug 2026
Agents

Can Coding Agents Solve Repository-Level Issues with Rendered Code? An Exploratory Study of Visual Representations

DGX agent

arXiv:2608.09268v1 Announce Type: cross Abstract: Visual modality has recently been explored as a way to compress textual tokens, including rendering code as images for static code understanding. We s

agentsarxiv-cs-ai
11 Aug 2026
Agents

CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems

DGX agent

arXiv:2608.09848v1 Announce Type: new Abstract: The development of embodied Intelligent Virtual Agents (IVAs) that have cognitive capabilities in real-time interactive virtual environments remains a c

agentsarxiv-cs-ai
11 Aug 2026
Local Ai

ColluSkill: Adversarial Cross-Skill Composition for Evading Agent Skill Scanners

DGX agent

arXiv:2608.09732v1 Announce Type: cross Abstract: Agent skills are emerging as an important attack surface in LLM-based agent systems. Through an empirical study of existing skill scanners, we find th

local-aiarxiv-cs-ai
11 Aug 2026
Agents

CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents

DGX agent

arXiv:2608.07855v1 Announce Type: new Abstract: Multi-turn Reasoning-and-Acting (ReAct) agents accumulate growing trajectories of reasoning, tool calls, and observations. Their key-value (KV) caches g

agentsarxiv-cs-lg
11 Aug 2026
Agents

Controlled Memory Interference in Continual LLM Agents

DGX agent

arXiv:2608.07622v1 Announce Type: new Abstract: Long-term memory enables AI agents to maintain continuity across sessions, personalize behavior, and evolve through accumulated experience. Yet memory e

agentsarxiv-cs-ai
11 Aug 2026
Agents

Improving Constraint Models with LLM Agents

DGX agent

arXiv:2608.08127v1 Announce Type: new Abstract: The runtime of Constraint Programming (CP) solvers is highly sensitive to modeling choices, such as symmetry breaking, implied constraints, global const

agentsarxiv-cs-ai
11 Aug 2026
Agents

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning

DGX agent

arXiv:2608.08255v1 Announce Type: cross Abstract: Agentic reinforcement learning (RL) often suffers from delayed and sparse rewards in real-world environments. A promising solution to this challenge i

agentsarxiv-cs-cl
11 Aug 2026
Agents

M^3Prune: Hierarchical Communication Graph Pruning for Efficient Multi-Modal Multi-Agent Retrieval-Augmented Generation

DGX agent

arXiv:2511.19969v2 Announce Type: replace Abstract: Recent advancements in multi-modal retrieval-augmented generation (mRAG), which enhance multi-modal large language models (MLLMs) with external know

agentsarxiv-cs-ai
11 Aug 2026
Agents

Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents

DGX agent

arXiv:2608.08389v1 Announce Type: new Abstract: Long-horizon research agents solve open-ended tasks through iterative retrieval, aggregation, and synthesis, but context grows rapidly while the margina

agentsarxiv-cs-ai
11 Aug 2026
Safety

Reflex First, Reflect Later: Latency-Aware Embodied LLM Agents for Dynamic Response

DGX agent

arXiv:2506.07223v2 Announce Type: replace Abstract: Large language models (LLMs) have substantially improved the planning capabilities of embodied agents, enabling their deployment in dynamic and safe

safetyarxiv-cs-ai
11 Aug 2026
Model Releases

SkillSentry: Reliable Skill Execution for LLM Agents via Runtime Assurance

DGX agent

arXiv:2608.09253v1 Announce Type: new Abstract: LLM agents are increasingly equipped with skills to perform complex tasks through multi-step reasoning and tool use. Although skills provide reusable pr

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

The Collaboration Gap: Exploration and Benchmarking of Open-World Agentic Cooperation

DGX agent

arXiv:2511.02687v2 Announce Type: replace Abstract: The trajectory of AI development suggests that we will increasingly rely on agent-based systems powered by language models, composed of independentl

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

The Scaffolding Matters More Than the Interface: A Controlled Comparison of MCP and CLI Tool Use Across Seven Agent Scaffoldings, Five Language Models, and One Software Task

DGX agent

arXiv:2608.08654v1 Announce Type: new Abstract: How much an AI coding agent costs to run can depend more on the agent scaffolding that drives it than on the interface through which it reaches its tool

model-releasesarxiv-cs-ai
11 Aug 2026
Agents

Agentic Planning for Symbolic Execution

DGX agent

arXiv:2608.06397v1 Announce Type: cross Abstract: Symbolic execution seeks to explore feasible program paths, yet a practical run may exhaust its resources while much program behaviour remains unreach

agentsarxiv-cs-ai
10 Aug 2026
Model Releases

Evaluating XAI Support From A Hierarchical Reinforcement Learning Policy in Human-Agent Collaboration

DGX agent

arXiv:2608.06381v1 Announce Type: cross Abstract: Explainable AI (XAI) has shown promise for human-agent collaboration, yet results rely on hand-crafted policies in custom environments, limiting gener

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

Long-Horizon Agent Trajectory Attribution: A Unified Benchmark and Fine-Grained Annotation Framework

DGX agent

arXiv:2608.06909v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly operate through long-horizon trajectories involving user instructions, tool use, external observations, a

model-releasesarxiv-cs-ai
10 Aug 2026
Local Ai

Online Monitoring and Corrective Steering of Programming Agents

DGX agent

arXiv:2608.06701v1 Announce Type: cross Abstract: Fixing GitHub issues in large-scale projects is a long-horizon task, especially when a fix requires changes across multiple locations or the issue des

local-aiarxiv-cs-ai
10 Aug 2026
Agents

AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games

DGX agent

arXiv:2608.06362v1 Announce Type: cross Abstract: Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or expert time.

agentsarxiv-cs-ai
7 Aug 2026
Safety

Breadcrumbing Search Agents

DGX agent

arXiv:2608.04565v1 Announce Type: cross Abstract: LLM-based search agents are widely used for information-seeking tasks, but their reliance on external tool returns introduces a critical security risk

safetyarxiv-cs-ai
6 Aug 2026
Safety

DAC-Pose: Dual-Agent Collaborative Framework for Pose-Guided Human Generation

DGX agent

arXiv:2608.04622v1 Announce Type: new Abstract: AI agents have emerged as a powerful new paradigm in generative image synthesis, enabling systems to perform complex semantic reasoning rather than pass

safetyarxiv-cs-cv
6 Aug 2026
Agents

EviGraph: Evidence-Guided Autonomous Research Agents

DGX agent

arXiv:2608.04738v1 Announce Type: new Abstract: Autonomous research agents can generate hypotheses, execute experiments, and draft manuscripts, yet their outputs often contain unsupported claims and i

agentsarxiv-cs-ai
6 Aug 2026
Model Releases

MatrAIx: Simulating the World with 8.3 Billion Persona Agents

DGX agent

arXiv:2608.04205v1 Announce Type: new Abstract: Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract aw

model-releasesarxiv-cs-ai
6 Aug 2026
Agents

OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents

DGX agent

arXiv:2608.05013v1 Announce Type: cross Abstract: LLM agents are increasingly applied to open-ended everyday requests that span work, study, and life. These tasks are long-horizon, cross-environment,

agentsarxiv-cs-ai
6 Aug 2026
Safety

PRIMAL3: Pathfinding via Reinforcement and Imitation Multi-Agent Learning - Leveraging LaCAM3

DGX agent

arXiv:2608.04905v1 Announce Type: new Abstract: We present PRIMAL3, an ultra-large-scale learning-based framework for multi-agent pathfinding (MAPF) that integrates reinforcement learning, topology-aw

safetyarxiv-cs-ro
6 Aug 2026
Model Releases

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks

DGX agent

arXiv:2608.03764v1 Announce Type: new Abstract: Agent self-evolution updates an agent's persistent state from prior experience and reuses it to solve related tasks more effectively. Evaluating self-ev

model-releasesarxiv-cs-ai
5 Aug 2026
Agents

Improving Sample Efficiency in Multi-Agent Reinforcement Learning for Simulated Football Games via Exploration

DGX agent

arXiv:2503.13077v2 Announce Type: replace Abstract: Multi-agent reinforcement learning has shown promise in learning cooperative behaviors in team-based environments. However, such methods often deman

agentsarxiv-cs-lg
5 Aug 2026
Agents

OR-Agent: Bridging Evolutionary Search and Structured Research for Automated Algorithm Discovery

DGX agent

arXiv:2602.13769v3 Announce Type: replace Abstract: Automating heuristic design in complex, experiment-driven domains requires more than iterative mutation of solution algorithms. Current LLM-based ev

agentsarxiv-cs-ai
5 Aug 2026
Agents

Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents

DGX agent

arXiv:2608.02751v1 Announce Type: cross Abstract: Existing deep-research agents use a search-visit workflow that retrieves and reads whole pages, without considering the addressable structure that web

agentsarxiv-cs-ai
5 Aug 2026
Agents

Training Documents Reranker with Search Rubrics for Deep Research Agent

DGX agent

arXiv:2608.03527v1 Announce Type: cross Abstract: Retrieval systems help deep research agents generate high-quality answers by providing relevant documents. However, existing retrievers typically sele

agentsarxiv-cs-ai
5 Aug 2026
Model Releases

Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking

DGX agent

arXiv:2606.15673v2 Announce Type: replace Abstract: Web agents act through long interaction sequences, yet existing benchmarks evaluate only terminal success, discarding all process information and of

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

AgentMemBench: A Systematic Benchmark for Evaluating Long-Term Memory Management Strategies in Conversational AI Agents

DGX agent

arXiv:2608.00009v1 Announce Type: new Abstract: Long-term memory remains a critical bottleneck for conversational AI agents, whose finite context windows cannot support coherent recall across thousand

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

DrawAI: Agentic Benchmark and Workflow for Making Raster Images Editable

DGX agent

arXiv:2608.00548v1 Announce Type: new Abstract: Recent image-generation models and multimodal agents can produce high-quality visuals for increasingly complex visual communication tasks. Yet their ras

model-releasesarxiv-cs-cv
4 Aug 2026
Agents

PGMem: Tightly Coupled Persona-Memory Graph for Lifelong Personalized Agents

DGX agent

arXiv:2608.01708v1 Announce Type: new Abstract: Long-term personalized dialogue agents must track user preferences as their personas evolve. Existing memory systems organize past events well, but stor

agentsarxiv-cs-cl
4 Aug 2026
Agents

Token-Native Storage: Read and Write in your Agent's Language

DGX agent

arXiv:2608.02376v1 Announce Type: cross Abstract: Search and database engines still store text as UTF-8, a format built for humans. But the systems that increasingly read and write that text (embedder

agentsarxiv-cs-cl
4 Aug 2026
Agents

Trajectories That Segment Themselves: Agent-Declared Boundaries as a Training Unit

DGX agent

arXiv:2608.02302v1 Announce Type: cross Abstract: Long-horizon coding-agent trajectories are poorly matched to the credit units available to train on: a single action has no stable value, an episode l

agentsarxiv-cs-lg
4 Aug 2026
Model Releases

AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-Correction

DGX agent

arXiv:2607.29549v1 Announce Type: new Abstract: Large language models have demonstrated strong mathematical problem-solving capabilities, yet reliably verifying their candidate answers remains challen

model-releasesarxiv-cs-ai
3 Aug 2026
Agents

EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration

DGX agent

arXiv:2607.18529v2 Announce Type: replace-cross Abstract: Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Exist

agentsarxiv-cs-ai
3 Aug 2026
Local Ai

RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems

DGX agent

arXiv:2607.29241v1 Announce Type: cross Abstract: Optimizing modern recommender models still depends heavily on engineers manually iterating over architectural, objective, and training-strategy change

local-aiarxiv-cs-ai
3 Aug 2026
Model Releases

SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition

DGX agent

arXiv:2607.28692v1 Announce Type: new Abstract: Large language model (LLM) agents have been increasingly adopted in scientific research for organizing and invoking specialized computational tools. How

model-releasesarxiv-cs-ai
3 Aug 2026
Safety

Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents

DGX agent

arXiv:2607.29254v1 Announce Type: new Abstract: AI agents extend large language models (LLMs) with external tools, enabling them to perform complex tasks and translate model outputs into consequential

safetyarxiv-cs-ai
3 Aug 2026
Model Releases

Validation Evidence in LLM Repair Agents: How Much of What Passes Actually Tests the Bug?

DGX agent

arXiv:2607.28871v1 Announce Type: cross Abstract: When a repair agent runs a test and sees it pass, the result is treated as evidence about the reported defect. We measure how often that treatment is

model-releasesarxiv-cs-ai
3 Aug 2026
Model Releases

A Graph-Native Bitemporal Memory Store for Conversational AI Agents

DGX agent

arXiv:2607.26520v1 Announce Type: cross Abstract: Conversational AI agents commonly lack persistent memory across sessions. The obvious fixes like injecting full chat histories into the context window

model-releasesarxiv-cs-ai
31 Jul 2026
Model Releases

ClinLens: Towards Long-Horizon Coding Agents for Longitudinal Multimodal Clinical Data Science

DGX agent

arXiv:2607.26155v1 Announce Type: new Abstract: Clinical data-science agents must transform heterogeneous longitudinal records into auditable analyses, yet existing benchmarks largely isolate medical

model-releasesarxiv-cs-ai
31 Jul 2026
Model Releases

Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale

DGX agent

arXiv:2607.28074v1 Announce Type: cross Abstract: Computer-use agents learn from what their actions change, so training one needs applications it can act on, break and reset. The applications that mat

model-releasesarxiv-cs-lg
31 Jul 2026
Model Releases

IDP AutoOpt: Agent-Driven Optimization of Document Processing Pipeline Configurations

DGX agent

arXiv:2607.26075v1 Announce Type: cross Abstract: We present IDP AutoOpt, an autonomous LLM agent that discovers high-performing configurations for intelligent document processing (IDP) pipelines. Tun

model-releasesarxiv-cs-ai
31 Jul 2026
← Previous
1…3637383940…233
Next →