AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
5,141 results
Local Ai

AgenticSCR: An Autonomous Agentic Secure Code Review for Immature Vulnerabilities Detection

DGX agent

arXiv:2601.19138v2 Announce Type: replace-cross Abstract: Secure code review is critical during pre-integration, where Atlassian developers rely on lightweight analysis tools, while deep security asse

local-aiarxiv-cs-ai
5 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Agentic Coding in the Wild: Characterizing GitHub Copilot Traces at Production Scale

DGX agent

arXiv:2608.00101v1 Announce Type: cross Abstract: AI coding agents like GitHub Copilot, Claude Code, and Codex interleave multi-step LLM inference with tool execution, creating a workload different fr

model-releasesarxiv-cs-lg
4 Aug 2026
Safety

CRISP: Critical Step Perception for Training Efficient Deep Search Agents

DGX agent

arXiv:2608.01867v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly extended into deep search agents that solve complex questions through multi-step interaction with external

safetyarxiv-cs-cl
4 Aug 2026
Model Releases

Prompt-Induced Waste in Large Reasoning Models: A Preregistered Two-Harness Benchmark of Coding Agents

DGX agent

arXiv:2608.01347v1 Announce Type: new Abstract: Large reasoning models used as coding agents incur costs from deliberation, tool calls, and repeated agent turns, yet the causal effect of prompt wordin

model-releasesarxiv-cs-cl
4 Aug 2026
Safety

Beyond Component Testing: Validating Agentic AI Systems

DGX agent

arXiv:2607.29405v1 Announce Type: new Abstract: Agentic AI systems act through multi-step trajectories that combine planning, tool use, memory, interaction, and adaptation. This behavior stretches val

safetyarxiv-cs-ai
3 Aug 2026
Agents

MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations

DGX agent

arXiv:2607.28956v1 Announce Type: new Abstract: Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success crite

agentsarxiv-cs-ai
3 Aug 2026
Safety

Unanticipated Effects of Generative AI on Expertise Pathways and Performance Perception in System Administration

DGX agent

arXiv:2607.28650v1 Announce Type: cross Abstract: While industry discourse often emphasizes immediate productivity gains and frames GenAI primarily as a tool for automation, the integration of GenAI i

safetyarxiv-cs-ai
3 Aug 2026
Agents

CRMWeaver: Building Powerful Business Agent via Agentic RL and Shared Memories

DGX agent

arXiv:2510.25333v2 Announce Type: replace Abstract: Recent years have witnessed the rapid development of LLM-based agents, which shed light on using language agents to solve complex real-world problem

agentsarxiv-cs-cl
31 Jul 2026
Safety

SimpleWikiSearch: A Clean Offline Wikipedia Environment for Agentic Search

DGX agent

arXiv:2607.26070v1 Announce Type: cross Abstract: Large language model (LLM)-based agentic search systems are often evaluated as if the underlying LLM were the only component that matters, yet their m

safetyarxiv-cs-ai
31 Jul 2026
Agents

Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability

DGX agent

arXiv:2607.26637v1 Announce Type: new Abstract: Deployed LLM agents increasingly keep their long-term memory as a filesystem: a directory tree of markdown files that the agent itself reads, writes, an

agentsarxiv-cs-cl
30 Jul 2026
Model Releases

Addressable Recall Compaction for Long Context-Window Control in AI Agents

DGX agent

arXiv:2607.25066v1 Announce Type: new Abstract: Long-horizon LLM agents accumulate reasoning traces, actions, and tool observations that can eventually exceed a model's fixed context window. Existing

model-releasesarxiv-cs-ai
29 Jul 2026
Safety

Context Assembly as the Controlled Variable: A Control-Theoretic View of Harness Policies for Frozen LLM Agents

DGX agent

arXiv:2607.25408v1 Announce Type: new Abstract: A growing body of 2026 work applies control theory to LLM agents: Lyapunov-certified stability for tool-mediated controllers (Prinos et al., 'Stable Age

safetyarxiv-cs-ai
29 Jul 2026
Agents

Are You Still the Agent I Authorized? Earned Authority under a Fixed Ceiling for Evolving Agents

DGX agent

arXiv:2607.23586v1 Announce Type: new Abstract: Long-lived AI agents increasingly evolve after deployment by retaining experience, acquiring skills and tools, revising workflows, delegating work, and

agentsarxiv-cs-ai
28 Jul 2026
Agents

CORVUS: Context Optimization and Reduction Via Underlying Synchronization for LLM Coding Agents

DGX agent

arXiv:2607.22711v1 Announce Type: cross Abstract: LLM coding agents operate by constructing trajectories that accumulate reasoning, tool calls, and results to enable multi-step decision-making. Howeve

agentsarxiv-cs-ai
28 Jul 2026
Safety

CRAFT: Learn the Schema, Execute the Plan

DGX agent

arXiv:2607.22642v1 Announce Type: new Abstract: Enterprise coding agents translate natural-language analytical requests into executable code over proprietary APIs, schemas, and metric definitions. Yet

safetyarxiv-cs-ai
28 Jul 2026
Research

From Vibe to Code -- and Back: Lexical Oscillation in the Formation of Design Intent with Generative AI

DGX agent

arXiv:2607.23126v1 Announce Type: cross Abstract: Generative AI design tools make natural-language prompts a starting point for design, placing new articulation demands on designers. Rather than treat

researcharxiv-cs-ai
28 Jul 2026
Agents

Plans Work in Mysterious Ways: Evaluating a Plan Mode for Spreadsheet Agents

DGX agent

arXiv:2607.23670v1 Announce Type: cross Abstract: Plan Modes have become standard features in agentic programming tools, allowing users to gain transparency and control by working with the agent to de

agentsarxiv-cs-ai
28 Jul 2026
Model Releases

Stability of AI Governance Systems: A Coupled Dynamics Model of Public Trust and Social Disruptions

DGX agent

arXiv:2603.20248v2 Announce Type: replace-cross Abstract: AI systems are increasingly entrenched in public governance, yet scholarship lacks formal tools to determine when deviations of public trust i

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions

DGX agent

arXiv:2607.21635v1 Announce Type: new Abstract: Personal agents maintain memories, learned skills, tool configurations, and policy state that evolve with each user. Existing agent benchmarks often eva

model-releasesarxiv-cs-lg
27 Jul 2026
Safety

The Ethics of Autonomous AI Agents for Offensive Security

DGX agent

arXiv:2607.20255v1 Announce Type: cross Abstract: LLM-driven autonomous agents are reshaping offensive security. Unlike traditional penetration-testing tooling -- deterministic, narrowly scoped, and o

safetyarxiv-cs-ai
23 Jul 2026
Model Releases

Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search

DGX agent

arXiv:2510.18939v2 Announce Type: replace Abstract: Long-horizon agentic search requires iteratively exploring the web over long trajectories and synthesizing information across many sources, enabling

model-releasesarxiv-cs-cl
15 Jul 2026
Model Releases

CausalDS: Benchmarking Causal Reasoning in Data-Science Agents

DGX agent

arXiv:2607.08093v1 Announce Type: new Abstract: Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use. Yet the relevant b

model-releasesarxiv-cs-ai
10 Jul 2026
Model Releases

aiAuthZ: Off-Host, Identity-Bound Authorization for AI Agents

DGX agent

arXiv:2607.05518v1 Announce Type: cross Abstract: AI agents issue tool calls on the basis of text they cannot verify, so any party who controls part of the context can forge the appearance of authorit

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents

DGX agent

arXiv:2607.06008v1 Announce Type: new Abstract: Large language model (LLM) agents have shown strong performance in long-horizon tasks that require planning, tool use, and interaction with external env

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

DualView: Preventing Indirect Prompt Injection in Personal AI Agents

DGX agent

arXiv:2607.03821v1 Announce Type: cross Abstract: Personal AI agents that run on the user's local machine, such as OpenClaw, automate daily tasks including web search, email, and file management. Thei

model-releasesarxiv-cs-ai
7 Jul 2026
Research

Causal Explanations for Image Classifiers

DGX agent

arXiv:2411.08875v4 Announce Type: replace Abstract: Existing algorithms for explaining the output of image classifiers use different definitions of explanations and a variety of techniques to find the

researcharxiv-cs-ai
3 Jul 2026
Model Releases

Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification

DGX agent

arXiv:2607.01793v1 Announce Type: new Abstract: LLM agents increasingly perform autonomous actions through external tools, leading to complex and evolving safety risks. However, existing safety testin

model-releasesarxiv-cs-ai
3 Jul 2026
Agents

Self-GC: Self-Governing Context for Long-Horizon LLM Agents

DGX agent

arXiv:2607.00692v1 Announce Type: new Abstract: Long-horizon LLM agents accumulate tool results, files, plans, and user constraints that are too structured to be treated as a disposable text suffix. C

agentsarxiv-cs-ai
2 Jul 2026
Model Releases

TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data?

DGX agent

arXiv:2606.13148v2 Announce Type: replace Abstract: Climate and environmental decision-making increasingly requires reasoning across heterogeneous inputs, including gridded physical data, satellite im

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

LLM Agents Are Latent Context Managers: Eliciting Self-Managed Context via a Proprioceptive Dashboard

DGX agent

arXiv:2606.30005v1 Announce Type: new Abstract: Long-horizon tool agents are bottlenecked by how their context grows toward the limits of the context window. Recent systems make context management age

model-releasesarxiv-cs-cl
30 Jun 2026
Research

wav2VOT: Automatic estimation of voice onset time, closure duration, and burst realisation with wav2vec2

DGX agent

arXiv:2606.28857v1 Announce Type: cross Abstract: While automatic tools for speech annotation are now commonplace within phonetic research pipelines, many tasks require substantial manual correction o

researcharxiv-cs-cl
30 Jun 2026
Safety

Radical AI Interpretability

DGX agent

arXiv:2606.26523v1 Announce Type: new Abstract: We develop a framework for interpreting AI systems as agents, drawing on the philosophical tradition of radical interpretation and the tools of mechanis

safetyarxiv-cs-ai
26 Jun 2026
Model Releases

BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents

DGX agent

arXiv:2605.06177v2 Announce Type: replace Abstract: Reproducing and comparing deep research agents today is hard: the same backbone evaluated on the same benchmark can report different accuracies acro

model-releasesarxiv-cs-ai
24 Jun 2026
Safety

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning

DGX agent

arXiv:2606.12195v1 Announce Type: new Abstract: Recent progress in foundation models has shifted toward agentic behavior involving multi-step reasoning and tool use. However, open-source efforts large

safetyarxiv-cs-cv
11 Jun 2026
Agents

ISE: An Execution-Grounded Recipe for Multi-Turn OS-Agent Trajectories

DGX agent

arXiv:2606.11520v1 Announce Type: cross Abstract: Training capable OS agents requires data that simultaneously captures structured user intents, multi-turn task delegation, and grounded tool execution

agentsarxiv-cs-ai
11 Jun 2026
Research

ros2probe: Non-intrusive, Kernel-selective Observability for Robot Operating System 2 Middleware

DGX agent

arXiv:2606.10746v1 Announce Type: new Abstract: Robot Operating System 2 (ROS 2), the de facto standard middleware framework for robots, runs each robot as a graph of nodes communicating over the Data

researcharxiv-cs-ro
10 Jun 2026
Agents

Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation

DGX agent

arXiv:2606.10749v1 Announce Type: cross Abstract: Large language model (LLM) agents are rapidly moving from conversational interfaces to software components that plan, invoke tools, maintain memory, a

agentsarxiv-cs-ai
10 Jun 2026
Safety

IEA: Amateur-Friendly Conversational Image Editing Agent via Three Stages of Multitask Alignment

DGX agent

arXiv:2606.08016v1 Announce Type: cross Abstract: Current image editing software often hinges on fixed filters or expert tuning, leaving a gap between amateur users' intent and outcomes. Creations by

safetyarxiv-cs-ai
9 Jun 2026
Safety

On-the-fly hand-eye calibration for the da Vinci surgical robot

DGX agent

arXiv:2601.14871v2 Announce Type: replace Abstract: In Robot-Assisted Minimally Invasive Surgery (RMIS), accurate tool localization is crucial to ensure patient safety and successful task execution. H

safetyarxiv-cs-ro
9 Jun 2026
Safety

PathoSage: Towards Multi-Source Evidence Adjudication in Pathology via Experience-Aware Agentic Workflow

DGX agent

arXiv:2606.07549v1 Announce Type: new Abstract: Recent advances in Multimodal Large Language Models (MLLMs) and agent workflows have shown strong promise for computational pathology, yet reliable patc

safetyarxiv-cs-ai
9 Jun 2026
Agents

The Token Not Taken: Sampling, State, and the Variability of AI Agent Outputs

DGX agent

arXiv:2606.08998v1 Announce Type: new Abstract: Agentic AI systems can behave differently across runs: the same request may produce a different plan, a different tool call, a different code edit, or a

agentsarxiv-cs-ai
9 Jun 2026
Model Releases

VATS: Exploiting Implicit Authority in Error-Path Injection via Systematic Mutation

DGX agent

arXiv:2606.07992v1 Announce Type: new Abstract: As the Model Context Protocol (MCP) standardizes tool-calling for autonomous agents, it introduces a critical, unexamined attack surface: the error-hand

model-releasesarxiv-cs-ai
9 Jun 2026
Agents

A Taxonomy of Runtime Faults in Model Context Protocol Servers

DGX agent

arXiv:2606.05339v1 Announce Type: cross Abstract: MCP (Model Context Protocol) enables LLMs (Large Language Models) to interact with external tools and data sources via a standardized protocol. Its ra

agentsarxiv-cs-ai
6 Jun 2026
Safety

Policy-Conditioned Counterfactual Credit for Verifiable Reinforcement Learning of Long-Horizon Language Agents

DGX agent

arXiv:2606.05263v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards improves reasoning and tool use, yet long-horizon language agents still learn unsupported evidence chai

safetyarxiv-cs-ai
6 Jun 2026
Local Ai

Framing Migration News with LLMs: Structured CoT as a Support for Human Interpretation

DGX agent

arXiv:2606.03761v1 Announce Type: new Abstract: Frame analysis of migration news is a socially consequential task: media scholars and researchers who study how migration is narrated need tools that ar

local-aiarxiv-cs-cl
3 Jun 2026
Model Releases

WRIT: Write-Read Intensive Trajectory Synthesis for Multi-Turn User-Facing Agents

DGX agent

arXiv:2606.02908v1 Announce Type: cross Abstract: Multi-turn user-facing agents must infer user intent from incomplete requests, collect missing information through dialogue and tools, and execute val

model-releasesarxiv-cs-ai
3 Jun 2026
Agents

Acting with AI: An Interaction-Based Framework for Agentic Tort Liability

DGX agent

arXiv:2606.00518v1 Announce Type: new Abstract: Agentic AI systems can plan over multiple steps, use tools, and execute tasks over time. When such systems cause harm, tort law struggles to allocate re

agentsarxiv-cs-ai
2 Jun 2026
Hardware

LayerRoute: Input-Conditioned Adaptive Layer Skipping via LoRA Fine-Tuning for Agentic Language Models

DGX agent

arXiv:2606.01838v1 Announce Type: cross Abstract: Agentic language model systems alternate between two structurally distinct step types: structured tool calls (short, deterministic, low perplexity) an

hardwarearxiv-cs-ai
2 Jun 2026
← Previous
1…1314151617…108
Next →