AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,289 results
Model Releases

VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning

DGX agent

arXiv:2607.02927v1 Announce Type: cross Abstract: Video understanding is moving beyond closed-context perception toward open-world evidence exploration, a paradigm formalized as Video Deep Research (V

model-releasesarxiv-cs-ai
7 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Adoption and Impact of Command-Line AI Coding Agents: A Study of Microsoft's Early 2026 Rollout of Claude Code and GitHub Copilot CLI

DGX agent

arXiv:2607.01418v1 Announce Type: cross Abstract: Organizations rolling out agentic command line tools like Anthropic's Claude Code and GitHub's Copilot CLI need to know who will try them, who will ke

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Agent4cs: A Multi-agent System for Code Summarization in Large Hierarchical Codebases

DGX agent

arXiv:2607.01425v1 Announce Type: new Abstract: Understanding large, complex codebases, especially those with obfuscated structures and incomplete documentation, remains a significant challenge. Exist

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

AgenticDataBench: A Comprehensive Benchmark for Data Agents

DGX agent

arXiv:2607.01647v1 Announce Type: cross Abstract: Data science aims to derive actionable insights from heterogeneous raw data, unlocking the value of the massive amounts of data generated in modern so

model-releasesarxiv-cs-ai
3 Jul 2026
Research

Multi-Head Recurrent Memory Agents

DGX agent

arXiv:2607.01523v1 Announce Type: cross Abstract: Recurrent memory agents extend LLMs to arbitrarily long contexts by iteratively consolidating input into a fixed-size memory window. Despite their sca

researcharxiv-cs-ai
3 Jul 2026
Safety

When Should Service Agents Reconsider? Difficulty-Routed Control in Customer-Service Operations

DGX agent

arXiv:2607.01426v1 Announce Type: new Abstract: Autonomous customer-service agents are shifting from conversational interfaces toward operational execution roles: they retrieve firm records, apply ser

safetyarxiv-cs-ai
3 Jul 2026
Model Releases

AGI Maze as a Benchmark Framework for World-Modeling Agents

DGX agent

arXiv:2607.00627v1 Announce Type: new Abstract: Large language models (LLMs) are powerful pattern-completion systems, but their default operating mode - predicting the next token from a static context

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?

DGX agent

arXiv:2607.01211v1 Announce Type: cross Abstract: Repository-level performance-optimization benchmarks such as GSO, SWE-Perf and SWE-fficiency evaluate coding agents by applying patches to real reposi

model-releasesarxiv-cs-ai
2 Jul 2026
Safety

Mnemosyne: Agentic Transaction Processing for Validating and Repairing AI-generated Workflows

DGX agent

arXiv:2607.00269v1 Announce Type: new Abstract: LLMs, solvers, and agent teams increasingly generate workflow actions, repairs, and plans, but a generated action may be syntactically valid yet stale,

safetyarxiv-cs-ai
2 Jul 2026
Local Ai

SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks

DGX agent

arXiv:2607.00053v1 Announce Type: cross Abstract: Large language models (LLMs) embedded in multi-turn agentic harnesses are reshaping software engineering (SWE), but routing every task to a frontier m

local-aiarxiv-cs-ai
2 Jul 2026
Agents

Urban Deceleration Behavior Modes Under Scene Context: An Early-Kinematic Classifier from Argoverse 2 Multi-Agent Trajectories

DGX agent

arXiv:2607.00027v1 Announce Type: cross Abstract: Urban deceleration is one of the most empirically studied yet least taxonomically organized behaviors in car-following research. Recent perception-equ

agentsarxiv-cs-lg
2 Jul 2026
Model Releases

A Semantic-Layer-Mediated Agent for Natural Language to SQL over Heterogeneous Enterprise Databases

DGX agent

arXiv:2606.31041v1 Announce Type: new Abstract: Natural language-to-SQL (NL2SQL) over real-world enterprise databases remains significantly more challenging than on academic benchmarks. Enterprise sch

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

An Executable Benchmarking Suite for Tool-Using Agents

DGX agent

arXiv:2605.11030v2 Announce Type: replace-cross Abstract: Closed-loop tool-using agents are increasingly evaluated in executable web, code, and micro-task environments, but benchmark reports often con

model-releasesarxiv-cs-ai
1 Jul 2026
Agents

DA-Studio: An Agentic System for End-to-End Data Analysis

DGX agent

arXiv:2606.31423v1 Announce Type: cross Abstract: Real-world data analysis is a multi-step process over heterogeneous inputs rather than merely producing a final answer. A practical system should auto

agentsarxiv-cs-ai
1 Jul 2026
Model Releases

The Decomposition Is the Fingerprint: Per-Component Identity for Agent Skills

DGX agent

arXiv:2606.31272v1 Announce Type: cross Abstract: AI agents increasingly acquire and execute skills at runtime: bundles of prompt instructions, executable code, and tool declarations fetched from mark

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

Think in English, Answer in Korean: Efficient Adaptation of Multilingual Tool-Using Agents

DGX agent

arXiv:2606.31648v1 Announce Type: new Abstract: We present LuckyStar 111B, a 111B-parameter hybrid reasoning model developed through a collaboration between Cohere and LG CNS for Korean-English enterp

model-releasesarxiv-cs-ai
1 Jul 2026
Safety

Agentic Tool Use in Large Language Models

DGX agent

arXiv:2604.00835v2 Announce Type: replace Abstract: Large language models are increasingly being deployed as autonomous agents yet their real world effectiveness depends on reliable tools for informat

safetyarxiv-cs-cl
30 Jun 2026
Model Releases

An Agentic AI Pipeline for Appliance-Level Energy Anomaly Detection and LLM-Driven Recommendations

DGX agent

arXiv:2606.28467v1 Announce Type: cross Abstract: Appliance-level energy monitoring in office buildings produces noisy alerts that non-expert facility managers struggle to use. This paper proposes an

model-releasesarxiv-cs-ai
30 Jun 2026
Safety

Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems

DGX agent

arXiv:2606.20470v2 Announce Type: replace-cross Abstract: Agentic AI systems increasingly rely on language-model components to interpret instructions, process external data, invoke tools, and coordina

safetyarxiv-cs-ai
30 Jun 2026
Agents

Boundary Degree as a Node-level Feature for Epidemic Scenario Identification in Agent-based Cascade Simulations

DGX agent

arXiv:2606.29596v1 Announce Type: cross Abstract: Characterizing the scenario underlying an epidemic from its disease cascade is an important task in simulation analytics. We propose boundary degree,

agentsarxiv-cs-lg
30 Jun 2026
Model Releases

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction

DGX agent

arXiv:2606.29445v1 Announce Type: cross Abstract: Video understanding is a fundamental capability for multimodal intelligence, and recent Multimodal Large Language Models (MLLMs) have achieved remarka

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios?

DGX agent

arXiv:2606.29920v1 Announce Type: new Abstract: Rubric-based scoring has become a widely used paradigm in model evaluation, typically with LLM-as-a-Judge (LaaJ) for rubric scoring. However, the reliab

model-releasesarxiv-cs-cl
30 Jun 2026
Safety

Characterizing Large Language Model Agentic Workflows: A Study on N8n Ecosystem

DGX agent

arXiv:2606.29116v1 Announce Type: new Abstract: Large Language Models (LLMs) are rapidly being adopted in low-code and no-code automation platforms, where non-expert users design workflows that combin

safetyarxiv-cs-ai
30 Jun 2026
Safety

CRAFT: Counterfactual Credit Assignment from Free Sibling Rollouts for Self-Distilled Agentic Reinforcement Learning

DGX agent

arXiv:2606.29476v1 Announce Type: cross Abstract: Self-distilled agentic reinforcement learning augments trajectory-level reward with a token-level distillation loss, using as its teacher the same pol

safetyarxiv-cs-ai
30 Jun 2026
Model Releases

DuoMem: Towards Capable On-Device Memory Agents via Dual-Space Distillation

DGX agent

arXiv:2606.29961v1 Announce Type: cross Abstract: Large Language Model (LLM)-based agents can solve complex procedural tasks by interacting with environments over multiple turns, but this ability typi

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

From Detecting Agency to Doing Work: Self-Caused Credit Builds a Durable Behavioral Self in a Minimal Spiking Agent

DGX agent

arXiv:2606.30191v1 Announce Type: new Abstract: How does an agent that can tell self from world come to be durably shaped by that distinction? Recent work shows that a predictive system can detect its

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

LEDGER: Scaling Agentic Document Editing with Dependency-aware Graph Retrieval

DGX agent

arXiv:2606.28379v1 Announce Type: cross Abstract: We introduce LEDGER to tackle the novel context engineering challenge of agentic document editing, where localized edits to long, structured documents

model-releasesarxiv-cs-ai
30 Jun 2026
Research

MedEvoEval: Evaluating Continual Evolution of Doctor Agents through Simulated Clinical Episodes

DGX agent

arXiv:2606.28900v1 Announce Type: new Abstract: Doctor agents are moving beyond single-turn answer generation toward evolving clinical decision systems. Within an outpatient episode, they acquire evid

researcharxiv-cs-ai
30 Jun 2026
Model Releases

MemDelta: Controlled Baselines and Hidden Confounds in Agent Memory Evaluation

DGX agent

arXiv:2606.29914v1 Announce Type: new Abstract: Agent memory systems are increasingly evaluated against RAG and full-context baselines, but reported gains often mix changes in the memory method with c

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation System

DGX agent

arXiv:2606.18112v3 Announce Type: replace-cross Abstract: Agentic navigation systems require a base navigation model whose observation strategy can be externally reconfigured at inference time, becaus

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

Selective Memory Retention for Long-Horizon LLM Agents

DGX agent

arXiv:2606.29178v1 Announce Type: new Abstract: When does retention matter for memory-augmented LLM agents? We study this with TraceRetain, a lightweight framework for bounded external memory in froze

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Autoformalization of Agent Instructions into Policy-as-Code

DGX agent

arXiv:2606.26649v1 Announce Type: new Abstract: Agent safety in high-stakes domains requires formal policy enforcement, but most existing approaches either rely on probabilistic guardrails (fine-tuned

model-releasesarxiv-cs-ai
26 Jun 2026
Safety

OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning

DGX agent

arXiv:2606.26790v1 Announce Type: new Abstract: Outcome-based reinforcement learning provides a stable optimization backbone for language agents, but its sparse trajectory-level rewards provide little

safetyarxiv-cs-cl
26 Jun 2026
Local Ai

Temporal Validity in Retrieval Memory: Eliminating Stale-Fact Errors for AI Agents over Evolving Knowledge

DGX agent

arXiv:2606.26511v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) gives agents access to accumulated knowledge, but has no model of time. When a fact changes (e.g., a function is

local-aiarxiv-cs-ai
26 Jun 2026
Model Releases

Agentic evolution of physically constrained foundation models

DGX agent

arXiv:2606.25532v1 Announce Type: cross Abstract: Artificial intelligence increasingly drives automated scientific discovery, yet contemporary generalist agents lack physical grounding, frequently hal

model-releasesarxiv-cs-lg
25 Jun 2026
Local Ai

BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents

DGX agent

arXiv:2606.25556v1 Announce Type: new Abstract: Stepwise group-based RL is an attractive way to train long-horizon LLM agents without a learned critic: it reuses multiple sampled rollouts to estimate

local-aiarxiv-cs-cl
25 Jun 2026
Model Releases

BrainAgent: A Large Language Model-Driven Multi-Agent Framework for Autonomous Brain Signal Understanding

DGX agent

arXiv:2606.25400v1 Announce Type: new Abstract: Brain-Computer Interfaces (BCIs) and brain signal understanding are pivotal for clinical health and next-generation interactions. Despite this significa

model-releasesarxiv-cs-ai
25 Jun 2026
Safety

Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents

DGX agent

arXiv:2606.26080v1 Announce Type: new Abstract: Process reward models enable fine-grained, step-level evaluation of LLMs, yet building them for agentic settings remains prohibitively difficult: long-h

safetyarxiv-cs-lg
25 Jun 2026
Model Releases

Staying In Character: Perspective-Bounded Memory For Book-Based Role-Playing Agents

DGX agent

arXiv:2606.25632v1 Announce Type: new Abstract: Recent LLM role-playing systems build character agents from novels by extracting characters, scenes, and relations. Yet long-narrative role-playing suff

model-releasesarxiv-cs-cl
25 Jun 2026
Research

TRUSTMEM: Learning Trustworthy Memory Consolidation for LLM Agents with Long-Term Memory

DGX agent

arXiv:2606.25161v1 Announce Type: new Abstract: Large language model (LLM) agents rely on long-term memory to support extended interactions and personalized assistance beyond finite context windows. E

researcharxiv-cs-ai
25 Jun 2026
Safety

LecturaAgents: A Multi-Agent Framework for Adaptive Personalized AI-Assisted Learning and Embodied Teaching

DGX agent

arXiv:2606.16428v2 Announce Type: replace-cross Abstract: Effective personalized AI-assisted learning demands systems that can not only generate accurate learner-specific educational materials, but al

safetyarxiv-cs-ai
24 Jun 2026
Safety

Multi-agent imitation learning with function approximation: Linear Markov games and beyond

DGX agent

arXiv:2602.22810v2 Announce Type: replace Abstract: In this work, we present the first theoretical analysis of multi-agent imitation learning (MAIL) in linear Markov games where both the transition dy

safetyarxiv-cs-lg
24 Jun 2026
Model Releases

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?

DGX agent

arXiv:2606.24530v1 Announce Type: new Abstract: We introduce NatureBench, a cross-discipline benchmark of 90 tasks distilled from peer-reviewed Nature-family publications, designed to evaluate whether

model-releasesarxiv-cs-cl
24 Jun 2026
Model Releases

AD-Bench: A Real-World, Trajectory-Aware Advertising Analytics Benchmark for LLM Agents

DGX agent

arXiv:2602.14257v2 Announce Type: replace-cross Abstract: While Large Language Model (LLM) agents have made remarkable progress on complex reasoning, evaluating them in real-world environments remains

model-releasesarxiv-cs-lg
23 Jun 2026
Agents

Compressing Observation History into Agent Memory: Distilling Transformers into Recurrent Transformers

DGX agent

arXiv:2606.21562v1 Announce Type: new Abstract: Transformers are AI's workhorse with strong performance in modeling sequential data, but their computational cost becomes prohibitive when processing lo

agentsarxiv-cs-cv
23 Jun 2026
Model Releases

ENVS: Environment-Native Verified Search for Long-Horizon GUI Agents

DGX agent

arXiv:2606.22948v1 Announce Type: cross Abstract: As multimodal agents move from interface understanding to real software control, successful trajectory discovery in live desktop environments becomes

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

GitOfThoughts: Version-Controlled Reasoning and Agent Memory You Can Replay, Diff, and Merge

DGX agent

arXiv:2606.14470v2 Announce Type: replace-cross Abstract: Large language model reasoning leaves no trace once it is done. The steps of a chain of thought disappear when the context window closes, a pr

model-releasesarxiv-cs-lg
23 Jun 2026
Safety

HoloAgent-0: A Unified Embodied Agent Framework with 3D Spatial Memory

DGX agent

arXiv:2606.23565v1 Announce Type: cross Abstract: LLM agents follow a practical execution loop in digital environments: they reason over structured states, invoke tools, inspect feedback, and revise a

safetyarxiv-cs-cv
23 Jun 2026
← Previous
1…8788899091…236
Next →