AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,153 results
Model Releases

VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments

DGX agent

arXiv:2506.02387v3 Announce Type: replace Abstract: Recent advancements in Vision Language Models (VLMs) have expanded their capabilities to interactive agent tasks, yet existing benchmarks remain lim

model-releasesarxiv-cs-ai
14 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Agentic Jackal: Live Execution and Semantic Value Grounding for Text-to-JQL

DGX agent

arXiv:2604.09470v1 Announce Type: new Abstract: Translating natural language into Jira Query Language (JQL) requires resolving ambiguous field references, instance-specific categorical values, and com

model-releasesarxiv-cs-cl
13 Apr 2026
Agents

Mina: A Multilingual LLM-Powered Legal Assistant Agent for Bangladesh for Empowering Access to Justice

DGX agent

arXiv:2511.08605v3 Announce Type: replace Abstract: Bangladesh's low-income population faces major barriers to affordable legal advice due to complex legal language, procedural opacity, and high costs

agentsarxiv-cs-cl
10 Apr 2026
Agents

RemoteAgent: Bridging Vague Human Intents and Earth Observation with RL-based Agentic MLLMs

DGX agent

arXiv:2604.07765v1 Announce Type: new Abstract: Earth Observation (EO) systems are essentially designed to support domain experts who often express their requirements through vague natural language ra

agentsarxiv-cs-cv
10 Apr 2026
Model Releases

Agent Safety Should Be a Runtime Contract

DGX agent

arXiv:2608.11274v1 Announce Type: cross Abstract: The dominant paradigm treats AI safety as a property to be instilled during model training via RLHF, DPO, or Constitutional AI. We argue this is struc

model-releasesarxiv-cs-ai
13 Aug 2026
Safety

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction

DGX agent

arXiv:2608.11772v1 Announce Type: new Abstract: Self-correction is particularly useful when a failure constrains the next repair. Coding agents benefit from this property because compilers, tests, and

safetyarxiv-cs-cl
13 Aug 2026
Model Releases

Self-Evolving Embodied Agents via Skill-Harness Evolution

DGX agent

arXiv:2608.11350v1 Announce Type: new Abstract: Embodied agents are increasingly built as systems around foundation models, where performance depends not only on model weights but also on the skills,

model-releasesarxiv-cs-cl
13 Aug 2026
Model Releases

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark

DGX agent

arXiv:2608.11469v1 Announce Type: cross Abstract: AI agents are rapidly improving in cybersecurity capabilities when the source code is available for analysis, yet much of the software most consequent

model-releasesarxiv-cs-ai
13 Aug 2026
Model Releases

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy

DGX agent

arXiv:2607.11175v2 Announce Type: replace Abstract: The growing ability of large language models and vision-language models to jointly interpret and reason over images and text is reshaping medical im

model-releasesarxiv-cs-ai
13 Aug 2026
Agents

Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems

DGX agent

arXiv:2608.11879v1 Announce Type: new Abstract: Long-running conversational agents increasingly rely on a memory system to avoid resending the whole conversation each turn, yet how much that costs to

agentsarxiv-cs-cl
13 Aug 2026
Safety

On The Statistical Limits of Self-Improving Agents

DGX agent

arXiv:2510.04399v3 Announce Type: replace Abstract: We develop a learning-theoretic framework for analyzing self-improving agents by decomposing self-modification into five axes. Within this framework

safetyarxiv-cs-ai
12 Aug 2026
Model Releases

Persistent Recursive Worlds Enable Autonomous Software Evolution

DGX agent

arXiv:2608.10450v1 Announce Type: cross Abstract: Complex software systems develop over timescales that exceed the lifespan of any individual coding agent. Most agentic software systems preserve conti

model-releasesarxiv-cs-ai
12 Aug 2026
Agents

Robust Multi-Agent Bandits with Heavy-Tailed Rewards and Information Asymmetry

DGX agent

arXiv:2608.10529v1 Announce Type: cross Abstract: The multi-armed bandit problem is a central framework in sequential decision-making, extensively studied under sub-Gaussian reward assumptions. Howeve

agentsarxiv-cs-ai
12 Aug 2026
Agents

SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models

DGX agent

arXiv:2608.10538v1 Announce Type: new Abstract: Agent skills represent a standardized format for packaging procedural knowledge and domain expertise, serving within agent harness systems as an essenti

agentsarxiv-cs-ai
12 Aug 2026
Model Releases

Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding

DGX agent

arXiv:2608.11095v1 Announce Type: new Abstract: Agentic coding READMEs like CLAUDE.md grow without bound in real repositories, stopping only when the repository retires or someone rewrites the file wh

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

AndroidReality: How Far Are Mobile Agents from the Real World?

DGX agent

arXiv:2608.07775v1 Announce Type: new Abstract: Mobile agents have achieved promising results on clean online benchmarks such as AndroidWorld, yet their performance often degrades sharply in real-worl

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception

DGX agent

arXiv:2608.08392v1 Announce Type: new Abstract: Large language models are increasingly deployed as autonomous agents that interact with the web through browsers. While recent progress has been driven

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Compiling and Benchmarking Task-State Horizons for Embodied Agents

DGX agent

arXiv:2608.08036v1 Announce Type: new Abstract: Frontier agentic models are increasingly deployed as high-level planners for long-horizon embodied tasks. Existing robotic benchmarks have advanced long

model-releasesarxiv-cs-ro
11 Aug 2026
Agents

LatticeMind: A Conflict-Aware Memory Primitive for Multi-Agent Systems

DGX agent

arXiv:2608.08236v1 Announce Type: new Abstract: Multi-agent LLM systems often fail not for lack of candidate answers, but because they have no persistent mechanism for deciding which incompatible clai

agentsarxiv-cs-ai
11 Aug 2026
Agents

Multi-agent discovery of practical quantum LDPC codes

DGX agent

arXiv:2608.08996v1 Announce Type: cross Abstract: Quantum low-density parity-check (qLDPC) codes can encode multiple logical qubits using sparse parity checks, yet searching for useful finite-length i

agentsarxiv-cs-ai
11 Aug 2026
Agents

SAKE: Structured Agentic Knowledge Extrapolation for Complex LLM Reasoning via Reinforcement Learning

DGX agent

arXiv:2505.15062v5 Announce Type: replace-cross Abstract: Knowledge extrapolation is the process of inferring novel information by combining and extending existing knowledge that is explicitly availab

agentsarxiv-cs-ai
11 Aug 2026
Safety

Software Engineering for and with GUI Agent

DGX agent

arXiv:2608.09278v1 Announce Type: cross Abstract: GUI agents have advanced rapidly, producing a growing body of frameworks, benchmarks, and applications. However, this growth has outpaced the maturity

safetyarxiv-cs-ai
11 Aug 2026
Agents

STAIR: Effective Incident Response Using an End-to-End Agentic Planning Framework

DGX agent

arXiv:2608.09524v1 Announce Type: cross Abstract: Incident response planning is critical for restoring compromised software systems after cyberattacks. Common practice relies on expert-driven playbook

agentsarxiv-cs-ai
11 Aug 2026
Model Releases

A Multi-Agent Framework for Automated Coarse-Grained Molecular Dynamics of Polymers

DGX agent

arXiv:2608.06694v1 Announce Type: new Abstract: Coarse-grained (CG) molecular dynamics extends polymer simulation beyond the scales accessible to all-atom (AA) methods, but bottom-up CG modeling is la

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

IB-RL: Isolated Bilateral Reinforcement Learning for Strategic Dialogue Agents

DGX agent

arXiv:2608.06735v1 Announce Type: new Abstract: Reinforcement learning (RL) has achieved strong results in improving large language models (LLMs) on tasks with stationary, verifiable rewards, such as

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

PHOENIX: Fine-Tuned SLM-Powered Autonomous Satellite Lifetime Extension via Predictive Self-Healing and Multi-Agent AI Recovery

DGX agent

arXiv:2608.07126v1 Announce Type: cross Abstract: Most CubeSats, small and low-cost satellites roughly the size of a shoebox, do not survive as long as they were designed to: a study of 178 missions f

model-releasesarxiv-cs-ai
10 Aug 2026
Agents

Plan-and-Avoid: Real-Time Aircraft Trajectory Coordination in a Multi-Agent Environment

DGX agent

arXiv:2608.06648v1 Announce Type: new Abstract: This paper presents a real-time Plan-and-Avoid (PAA framework for coordinating cooperative multi-agent airspace operations around a declared priority tr

agentsarxiv-cs-ro
10 Aug 2026
Model Releases

Same physical state, different collective dynamics: state encodings select synchronization outcomes in language-model agents

DGX agent

arXiv:2608.06968v1 Announce Type: cross Abstract: Language-model agents act on state encodings of their environment, yet these are treated as interchangeable interfaces. Using pretrained language mode

model-releasesarxiv-cs-ai
10 Aug 2026
Safety

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents

DGX agent

arXiv:2608.06663v1 Announce Type: new Abstract: Frontier language models solve reasoning problems in a single forward pass that would have been research contributions years ago, yet fail at multi-hour

safetyarxiv-cs-cl
10 Aug 2026
Agents

Comparative Approaches to Agent Retrieval over Large Skill Libraries

DGX agent

arXiv:2608.06196v1 Announce Type: new Abstract: Agents backed by large skill libraries must decide which skills to load and in what order. Loading the entire library into context is expensive and prov

agentsarxiv-cs-ai
7 Aug 2026
Agents

Post-Hoc Trajectory-Risk Certification for Modular LLM-Based Security Agents

DGX agent

arXiv:2608.05199v1 Announce Type: cross Abstract: Autonomous security agents operate as staged pipelines, such as classifying network traffic and then attributing attacks to a specific technique. Spli

agentsarxiv-cs-ai
7 Aug 2026
Model Releases

StepReflect: Structured UI Transition Reflection for Mobile GUI Agents

DGX agent

arXiv:2608.05587v1 Announce Type: new Abstract: Autonomous mobile GUI agents require accurate action reflection for reliable long-horizon execution. Existing approaches rely on open-ended multimodal r

model-releasesarxiv-cs-ai
7 Aug 2026
Agents

WorldClaw: Agentic 3D Open-World Generation at Scale

DGX agent

arXiv:2608.05248v1 Announce Type: new Abstract: Generating large-scale, freely explorable 3D worlds from open-ended text remains challenging because a system must jointly maintain global spatial coher

agentsarxiv-cs-ai
7 Aug 2026
Research

ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment

DGX agent

arXiv:2608.05102v1 Announce Type: new Abstract: Long-horizon search agents must make multiple sequential actions (steps) to search, retrieve, verify, and integrate evidence to reach a final answer. Ho

researcharxiv-cs-ai
6 Aug 2026
Agents

Behavioral Skill Reconstruction: Reconstructing Hidden Functionality from LLM Agent Skills

DGX agent

arXiv:2608.04192v1 Announce Type: cross Abstract: Closed source agent skills may encode proprietary instructions, scripts, constants, and data. Providers may offer their capabilities as services while

agentsarxiv-cs-ai
6 Aug 2026
Agents

Embedding Large Language Models into Flow Controls: An Agentic Framework for Adaptive and Trustworthy Automated Cooking

DGX agent

arXiv:2608.04768v1 Announce Type: new Abstract: Automated cooking robots have traditionally relied on predefined procedures and rule-based control, ensuring stable execution but offering limited perso

agentsarxiv-cs-cv
6 Aug 2026
Model Releases

InsightEmb: Learning Action-Intent Embeddings for Agentic Insight Retrieval

DGX agent

arXiv:2608.04761v1 Announce Type: cross Abstract: Self-improving agents accumulate reusable insights from prior trajectories, making retrieval increasingly important for turning accumulated experience

model-releasesarxiv-cs-ai
6 Aug 2026
Model Releases

Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?

DGX agent

arXiv:2608.04828v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on skills, structured documents that specify when to act, which procedure to follow, and which tools

model-releasesarxiv-cs-cl
6 Aug 2026
Safety

ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation

DGX agent

arXiv:2608.04436v1 Announce Type: new Abstract: Text-to-image (T2I) models can produce visually compelling images, yet they remain limited on open-world tasks that require complex semantic understandi

safetyarxiv-cs-cv
6 Aug 2026
Model Releases

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)

DGX agent

arXiv:2608.04317v1 Announce Type: cross Abstract: Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evaluated almost

model-releasesarxiv-cs-ai
6 Aug 2026
Safety

When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents

DGX agent

arXiv:2608.04574v1 Announce Type: new Abstract: Memory-augmented VLM agents act on persistent spatial knowledge, yet that knowledge silently goes stale as the environment changes. We ask what happens

safetyarxiv-cs-cl
6 Aug 2026
Agents

ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?

DGX agent

arXiv:2608.03874v1 Announce Type: new Abstract: Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear whether these syst

agentsarxiv-cs-ai
5 Aug 2026
Agents

ETA: A New Agentic Paradigm for Embodied Tasks

DGX agent

arXiv:2608.03924v1 Announce Type: new Abstract: When will robots have their ChatGPT moment? Such a breakthrough requires a general-purpose robot that can handle unfamiliar tasks in unfamiliar environm

agentsarxiv-cs-ro
5 Aug 2026
Agents

Field Aware Agent Skill Retrieval

DGX agent

arXiv:2608.02880v1 Announce Type: cross Abstract: As lifelong learning agents accumulate lifelong growing skill banks, retrieving the correct skill becomes an increasingly important bottleneck. Most c

agentsarxiv-cs-lg
5 Aug 2026
Safety

Permission Denied: Policy-Graded Evaluation of Coding Agents in Hardened Environments

DGX agent

arXiv:2608.02670v1 Announce Type: cross Abstract: Coding agents increasingly run inside organizations whose security controls (scoped credentials, restricted egress, read-only filesystems, non-root ex

safetyarxiv-cs-ai
5 Aug 2026
Local Ai

SKILL-KD: Contrastive Skill Distillation for LLM Agents

DGX agent

arXiv:2607.28048v2 Announce Type: replace Abstract: Skill-based prompting has become a practical mechanism for improving large language model (LLM) agents, yet existing skill acquisition methods often

local-aiarxiv-cs-ai
5 Aug 2026
Model Releases

AdvPlan-Bench: Adversarial Evaluation of Structured Plan-Generation Agents

DGX agent

arXiv:2608.00832v1 Announce Type: new Abstract: Structured plan-generation agents are often evaluated as if a plan has quality in isolation, yet many realistic planning tasks require asking how a cand

model-releasesarxiv-cs-lg
4 Aug 2026
Model Releases

Agentic Coding in the Wild: Characterizing GitHub Copilot Traces at Production Scale

DGX agent

arXiv:2608.00101v1 Announce Type: cross Abstract: AI coding agents like GitHub Copilot, Claude Code, and Codex interleave multi-step LLM inference with tool execution, creating a workload different fr

model-releasesarxiv-cs-lg
4 Aug 2026
← Previous
1…5354555657…233
Next →