AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,950 results
Safety

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction

DGX agent

arXiv:2608.11772v1 Announce Type: new Abstract: Self-correction is particularly useful when a failure constrains the next repair. Coding agents benefit from this property because compilers, tests, and

safetyarxiv-cs-cl
13 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Agents

our outbound harness is focused on delivering the best economics we can, we're continuing to push on how our agents work to get another 90%+…

DGX agent

our outbound harness is focused on delivering the best economics we can, we're continuing to push on how our agents work to get another 90%+ cost savings for customers great talk by @hwchase17 and @He

agentsharrison-chase--x
13 Aug 2026
Model Releases

Qdrant and Minima Deliver 2.92x More Agentic RAG Tasks per GPU-Hour

DGX agent

Reducing Retrieval and Calls When a retrieval-augmented generation (RAG) agent runs, it often has to plan a search, check the evidence it gets back, and try again when that evidence falls short. Those

model-releasesqdrant
13 Aug 2026
Model Releases

Self-Evolving Embodied Agents via Skill-Harness Evolution

DGX agent

arXiv:2608.11350v1 Announce Type: new Abstract: Embodied agents are increasingly built as systems around foundation models, where performance depends not only on model weights but also on the skills,

model-releasesarxiv-cs-cl
13 Aug 2026
Model Releases

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark

DGX agent

arXiv:2608.11469v1 Announce Type: cross Abstract: AI agents are rapidly improving in cybersecurity capabilities when the source code is available for analysis, yet much of the software most consequent

model-releasesarxiv-cs-ai
13 Aug 2026
Model Releases

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy

DGX agent

arXiv:2607.11175v2 Announce Type: replace Abstract: The growing ability of large language models and vision-language models to jointly interpret and reason over images and text is reshaping medical im

model-releasesarxiv-cs-ai
13 Aug 2026
Agents

Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems

DGX agent

arXiv:2608.11879v1 Announce Type: new Abstract: Long-running conversational agents increasingly rely on a memory system to avoid resending the whole conversation each turn, yet how much that costs to

agentsarxiv-cs-cl
13 Aug 2026
Safety

On The Statistical Limits of Self-Improving Agents

DGX agent

arXiv:2510.04399v3 Announce Type: replace Abstract: We develop a learning-theoretic framework for analyzing self-improving agents by decomposing self-modification into five axes. Within this framework

safetyarxiv-cs-ai
12 Aug 2026
Model Releases

Persistent Recursive Worlds Enable Autonomous Software Evolution

DGX agent

arXiv:2608.10450v1 Announce Type: cross Abstract: Complex software systems develop over timescales that exceed the lifespan of any individual coding agent. Most agentic software systems preserve conti

model-releasesarxiv-cs-ai
12 Aug 2026
Agents

Robust Multi-Agent Bandits with Heavy-Tailed Rewards and Information Asymmetry

DGX agent

arXiv:2608.10529v1 Announce Type: cross Abstract: The multi-armed bandit problem is a central framework in sequential decision-making, extensively studied under sub-Gaussian reward assumptions. Howeve

agentsarxiv-cs-ai
12 Aug 2026
Agents

SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models

DGX agent

arXiv:2608.10538v1 Announce Type: new Abstract: Agent skills represent a standardized format for packaging procedural knowledge and domain expertise, serving within agent harness systems as an essenti

agentsarxiv-cs-ai
12 Aug 2026
Model Releases

Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding

DGX agent

arXiv:2608.11095v1 Announce Type: new Abstract: Agentic coding READMEs like CLAUDE.md grow without bound in real repositories, stopping only when the repository retires or someone rewrites the file wh

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

AndroidReality: How Far Are Mobile Agents from the Real World?

DGX agent

arXiv:2608.07775v1 Announce Type: new Abstract: Mobile agents have achieved promising results on clean online benchmarks such as AndroidWorld, yet their performance often degrades sharply in real-worl

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception

DGX agent

arXiv:2608.08392v1 Announce Type: new Abstract: Large language models are increasingly deployed as autonomous agents that interact with the web through browsers. While recent progress has been driven

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Compiling and Benchmarking Task-State Horizons for Embodied Agents

DGX agent

arXiv:2608.08036v1 Announce Type: new Abstract: Frontier agentic models are increasingly deployed as high-level planners for long-horizon embodied tasks. Existing robotic benchmarks have advanced long

model-releasesarxiv-cs-ro
11 Aug 2026
Agents

LatticeMind: A Conflict-Aware Memory Primitive for Multi-Agent Systems

DGX agent

arXiv:2608.08236v1 Announce Type: new Abstract: Multi-agent LLM systems often fail not for lack of candidate answers, but because they have no persistent mechanism for deciding which incompatible clai

agentsarxiv-cs-ai
11 Aug 2026
Agents

Multi-agent discovery of practical quantum LDPC codes

DGX agent

arXiv:2608.08996v1 Announce Type: cross Abstract: Quantum low-density parity-check (qLDPC) codes can encode multiple logical qubits using sparse parity checks, yet searching for useful finite-length i

agentsarxiv-cs-ai
11 Aug 2026
Agents

SAKE: Structured Agentic Knowledge Extrapolation for Complex LLM Reasoning via Reinforcement Learning

DGX agent

arXiv:2505.15062v5 Announce Type: replace-cross Abstract: Knowledge extrapolation is the process of inferring novel information by combining and extending existing knowledge that is explicitly availab

agentsarxiv-cs-ai
11 Aug 2026
Safety

Software Engineering for and with GUI Agent

DGX agent

arXiv:2608.09278v1 Announce Type: cross Abstract: GUI agents have advanced rapidly, producing a growing body of frameworks, benchmarks, and applications. However, this growth has outpaced the maturity

safetyarxiv-cs-ai
11 Aug 2026
Agents

SpaceXAI rolls out Grok Bot AI agent app in beta on Mac, iOS, Windows, and Linux, initially for SuperGrok Heavy, Cursor Ultra, and Cursor Teams Premium users (Zac Hall/9to5Mac)

DGX agent

Zac Hall / 9to5Mac: SpaceXAI rolls out Grok Bot AI agent app in beta on Mac, iOS, Windows, and Linux, initially for SuperGrok Heavy, Cursor Ultra, and Cursor Teams Premium users — SpaceXAI and Cursor

agentstechmeme
11 Aug 2026
Agents

STAIR: Effective Incident Response Using an End-to-End Agentic Planning Framework

DGX agent

arXiv:2608.09524v1 Announce Type: cross Abstract: Incident response planning is critical for restoring compromised software systems after cyberattacks. Common practice relies on expert-driven playbook

agentsarxiv-cs-ai
11 Aug 2026
Model Releases

A Multi-Agent Framework for Automated Coarse-Grained Molecular Dynamics of Polymers

DGX agent

arXiv:2608.06694v1 Announce Type: new Abstract: Coarse-grained (CG) molecular dynamics extends polymer simulation beyond the scales accessible to all-atom (AA) methods, but bottom-up CG modeling is la

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

IB-RL: Isolated Bilateral Reinforcement Learning for Strategic Dialogue Agents

DGX agent

arXiv:2608.06735v1 Announce Type: new Abstract: Reinforcement learning (RL) has achieved strong results in improving large language models (LLMs) on tasks with stationary, verifiable rewards, such as

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

PHOENIX: Fine-Tuned SLM-Powered Autonomous Satellite Lifetime Extension via Predictive Self-Healing and Multi-Agent AI Recovery

DGX agent

arXiv:2608.07126v1 Announce Type: cross Abstract: Most CubeSats, small and low-cost satellites roughly the size of a shoebox, do not survive as long as they were designed to: a study of 178 missions f

model-releasesarxiv-cs-ai
10 Aug 2026
Agents

Plan-and-Avoid: Real-Time Aircraft Trajectory Coordination in a Multi-Agent Environment

DGX agent

arXiv:2608.06648v1 Announce Type: new Abstract: This paper presents a real-time Plan-and-Avoid (PAA framework for coordinating cooperative multi-agent airspace operations around a declared priority tr

agentsarxiv-cs-ro
10 Aug 2026
Model Releases

Same physical state, different collective dynamics: state encodings select synchronization outcomes in language-model agents

DGX agent

arXiv:2608.06968v1 Announce Type: cross Abstract: Language-model agents act on state encodings of their environment, yet these are treated as interchangeable interfaces. Using pretrained language mode

model-releasesarxiv-cs-ai
10 Aug 2026
Safety

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents

DGX agent

arXiv:2608.06663v1 Announce Type: new Abstract: Frontier language models solve reasoning problems in a single forward pass that would have been research contributions years ago, yet fail at multi-hour

safetyarxiv-cs-cl
10 Aug 2026
Agents

Together Serverless Inference gives developers a high-throughput, managed production path for running Muse Glimmer across agentic and multim…

DGX agent

Together Serverless Inference gives developers a high-throughput, managed production path for running Muse Glimmer across agentic and multimodal workloads. Start building: https://www.together.ai/mode

agentstogether-ai--x
10 Aug 2026
Agents

Comparative Approaches to Agent Retrieval over Large Skill Libraries

DGX agent

arXiv:2608.06196v1 Announce Type: new Abstract: Agents backed by large skill libraries must decide which skills to load and in what order. Loading the entire library into context is expensive and prov

agentsarxiv-cs-ai
7 Aug 2026
Agents

Post-Hoc Trajectory-Risk Certification for Modular LLM-Based Security Agents

DGX agent

arXiv:2608.05199v1 Announce Type: cross Abstract: Autonomous security agents operate as staged pipelines, such as classifying network traffic and then attributing attacks to a specific technique. Spli

agentsarxiv-cs-ai
7 Aug 2026
Model Releases

StepReflect: Structured UI Transition Reflection for Mobile GUI Agents

DGX agent

arXiv:2608.05587v1 Announce Type: new Abstract: Autonomous mobile GUI agents require accurate action reflection for reliable long-horizon execution. Existing approaches rely on open-ended multimodal r

model-releasesarxiv-cs-ai
7 Aug 2026
Agents

This might be a good time to mention my recent keynotes titled: 'Agentic AI is Neurosymbolic AI'

DGX agent

This might be a good time to mention my recent keynotes titled: 'Agentic AI is Neurosymbolic AI' I would have assumed it was fairly obvious, but in case it's not: a million-line codebase (also known a

agentsfrancois-chollet--x
7 Aug 2026
Agents

WorldClaw: Agentic 3D Open-World Generation at Scale

DGX agent

arXiv:2608.05248v1 Announce Type: new Abstract: Generating large-scale, freely explorable 3D worlds from open-ended text remains challenging because a system must jointly maintain global spatial coher

agentsarxiv-cs-ai
7 Aug 2026
Research

ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment

DGX agent

arXiv:2608.05102v1 Announce Type: new Abstract: Long-horizon search agents must make multiple sequential actions (steps) to search, retrieve, verify, and integrate evidence to reach a final answer. Ho

researcharxiv-cs-ai
6 Aug 2026
Agents

Behavioral Skill Reconstruction: Reconstructing Hidden Functionality from LLM Agent Skills

DGX agent

arXiv:2608.04192v1 Announce Type: cross Abstract: Closed source agent skills may encode proprietary instructions, scripts, constants, and data. Providers may offer their capabilities as services while

agentsarxiv-cs-ai
6 Aug 2026
Agents

Embedding Large Language Models into Flow Controls: An Agentic Framework for Adaptive and Trustworthy Automated Cooking

DGX agent

arXiv:2608.04768v1 Announce Type: new Abstract: Automated cooking robots have traditionally relied on predefined procedures and rule-based control, ensuring stable execution but offering limited perso

agentsarxiv-cs-cv
6 Aug 2026
Model Releases

InsightEmb: Learning Action-Intent Embeddings for Agentic Insight Retrieval

DGX agent

arXiv:2608.04761v1 Announce Type: cross Abstract: Self-improving agents accumulate reusable insights from prior trajectories, making retrieval increasingly important for turning accumulated experience

model-releasesarxiv-cs-ai
6 Aug 2026
Model Releases

OpenAI debuts Agent Plugins, an open standard for bundling skills and MCP servers, and says Amazon, Cursor, Microsoft, and Vercel are on its steering committee (Zac Hall/9to5Mac)

DGX agent

Zac Hall / 9to5Mac: OpenAI debuts Agent Plugins, an open standard for bundling skills and MCP servers, and says Amazon, Cursor, Microsoft, and Vercel are on its steering committee — OpenAI's GPT-5 tur

model-releasestechmeme
6 Aug 2026
Model Releases

Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?

DGX agent

arXiv:2608.04828v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on skills, structured documents that specify when to act, which procedure to follow, and which tools

model-releasesarxiv-cs-cl
6 Aug 2026
Safety

ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation

DGX agent

arXiv:2608.04436v1 Announce Type: new Abstract: Text-to-image (T2I) models can produce visually compelling images, yet they remain limited on open-world tasks that require complex semantic understandi

safetyarxiv-cs-cv
6 Aug 2026
Model Releases

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)

DGX agent

arXiv:2608.04317v1 Announce Type: cross Abstract: Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evaluated almost

model-releasesarxiv-cs-ai
6 Aug 2026
Safety

When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents

DGX agent

arXiv:2608.04574v1 Announce Type: new Abstract: Memory-augmented VLM agents act on persistent spatial knowledge, yet that knowledge silently goes stale as the environment changes. We ask what happens

safetyarxiv-cs-cl
6 Aug 2026
Agents

ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?

DGX agent

arXiv:2608.03874v1 Announce Type: new Abstract: Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear whether these syst

agentsarxiv-cs-ai
5 Aug 2026
Agents

ETA: A New Agentic Paradigm for Embodied Tasks

DGX agent

arXiv:2608.03924v1 Announce Type: new Abstract: When will robots have their ChatGPT moment? Such a breakthrough requires a general-purpose robot that can handle unfamiliar tasks in unfamiliar environm

agentsarxiv-cs-ro
5 Aug 2026
Agents

Field Aware Agent Skill Retrieval

DGX agent

arXiv:2608.02880v1 Announce Type: cross Abstract: As lifelong learning agents accumulate lifelong growing skill banks, retrieving the correct skill becomes an increasingly important bottleneck. Most c

agentsarxiv-cs-lg
5 Aug 2026
Safety

Permission Denied: Policy-Graded Evaluation of Coding Agents in Hardened Environments

DGX agent

arXiv:2608.02670v1 Announce Type: cross Abstract: Coding agents increasingly run inside organizations whose security controls (scoped credentials, restricted egress, read-only filesystems, non-root ex

safetyarxiv-cs-ai
5 Aug 2026
Local Ai

SKILL-KD: Contrastive Skill Distillation for LLM Agents

DGX agent

arXiv:2607.28048v2 Announce Type: replace Abstract: Skill-based prompting has become a practical mechanism for improving large language model (LLM) agents, yet existing skill acquisition methods often

local-aiarxiv-cs-ai
5 Aug 2026
Model Releases

AdvPlan-Bench: Adversarial Evaluation of Structured Plan-Generation Agents

DGX agent

arXiv:2608.00832v1 Announce Type: new Abstract: Structured plan-generation agents are often evaluated as if a plan has quality in isolation, yet many realistic planning tasks require asking how a cand

model-releasesarxiv-cs-lg
4 Aug 2026
← Previous
1…9596979899…374
Next →