AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
5,141 results
Safety

ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents

DGX agent

arXiv:2606.31392v1 Announce Type: new Abstract: Tool-augmented vision-language models (VLMs) can solve multimodal, multi-step tasks by calling external tools, yet they remain fragile in practice. Exis

safetyarxiv-cs-ai
1 Jul 2026
Agents
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

VTOS: Learning to Orchestrate Vision Tools by Co-Searching Solutions and Observers

DGX agent

arXiv:2606.20728v1 Announce Type: new Abstract: Vision foundation tools such as open-vocabulary detectors, segmentation models, and post-processing operators are powerful building blocks for computer

agentsarxiv-cs-cv
23 Jun 2026
Safety

IAPO: Input Attribution-Aware Policy Optimization for Tool Use in Small Multimodal Agents

DGX agent

arXiv:2606.11652v1 Announce Type: new Abstract: This paper investigates reinforcement learning (RL) methods for improving tool-calling capabilities in multimodal small language model (SLM) agents. Whi

safetyarxiv-cs-lg
11 Jun 2026
Agents

Pushing the Limits of LLM Tool Calling via Experiential Knowledge Integration and Activation

DGX agent

arXiv:2606.10875v1 Announce Type: new Abstract: Large language models (LLMs) rely on tool use to act as autonomous agents, yet often fail in multi-step execution due to insufficient tool-related knowl

agentsarxiv-cs-cl
10 Jun 2026
Safety

Capability-Aligned Hierarchical Learning for Tool-Augmented LLMs

DGX agent

arXiv:2606.09371v1 Announce Type: new Abstract: Tool learning enables LLMs to invoke external tools to accomplish tasks. Prior studies have demonstrated the effectiveness of a hierarchical structure:

safetyarxiv-cs-ai
9 Jun 2026
Agents

Synthesize and Reward -- Reinforcement Learning for Multi-Step Tool Use in Live Environments

DGX agent

arXiv:2606.03892v1 Announce Type: cross Abstract: Training LLMs to orchestrate multi-step tool calls is held back by three coupled obstacles: realistic stateful execution environments are costly to bu

agentsarxiv-cs-ai
3 Jun 2026
Model Releases

Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation

DGX agent

arXiv:2512.16310v3 Announce Type: replace-cross Abstract: LLM-based agents increasingly use multiple external tools to complete complex tasks. We study Tools Orchestration Privacy Risk (TOP-R): an age

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

LogDx-CI: Benchmarking Log Reduction Tools for LLM Root-Cause Diagnosis

DGX agent

arXiv:2605.28876v1 Announce Type: cross Abstract: CI failure logs are large (median 5k lines, max 200k in this corpus) and noisy. Coding agents that try to debug them depend on an upstream tool to red

model-releasesarxiv-cs-ai
29 May 2026
Applications

Augment Engineering: A Methodology for Multi-Tool AI Orchestration Across Professional Domains

DGX agent

arXiv:2605.26146v1 Announce Type: cross Abstract: Organizations increasingly deploy separate purpose-built AI tools across professional domains, often hiring domain specialists for each, recreating th

applicationsarxiv-cs-ai
27 May 2026
Safety

ChainCaps: Composition-Safe Tool-Using Agents via Monotonic Capability Attenuation

DGX agent

arXiv:2605.26542v1 Announce Type: cross Abstract: Tool-using agents increasingly operate in open-ended deployment environments, where they compose file systems, web APIs, code interpreters, and enterp

safetyarxiv-cs-ai
27 May 2026
Agents

Tool-Call Dependency Structure is Linearly Decodable in LLM Agent Residual Streams

DGX agent

arXiv:2605.25310v1 Announce Type: new Abstract: Tool-using LLM agents produce trajectories whose calls form a directed dependency graph: earlier tool outputs supply arguments to later calls. Whether t

agentsarxiv-cs-cl
26 May 2026
Model Releases

Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIs

DGX agent

arXiv:2605.17558v1 Announce Type: cross Abstract: Training tool-calling agents requires large-scale trajectory data with verifiable labels, yet existing approaches either synthesize environments that

model-releasesarxiv-cs-cl
19 May 2026
Safety

PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools

DGX agent

arXiv:2604.01532v2 Announce Type: replace Abstract: LLM agents are beginning to invoke industrial asset-management tools through the Model Context Protocol (MCP), yet whether they can act reliably on

safetyarxiv-cs-ai
12 May 2026
Research

Useful for Exploration, Risky for Precision: Evaluating AI Tools in Academic Research

DGX agent

arXiv:2605.10125v1 Announce Type: new Abstract: Artificial intelligence (AI) tools are being incorporated into scientific research workflows with the potential to enhance efficiency in tasks such as d

researcharxiv-cs-ai
12 May 2026
Model Releases

AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents

DGX agent

arXiv:2605.07926v1 Announce Type: new Abstract: As LLM-based agents increasingly rely on external tools, it is important to evaluate their ability to sustain tool-grounded reasoning beyond familiar wo

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers

DGX agent

arXiv:2602.00933v2 Announce Type: replace-cross Abstract: The Model Context Protocol (MCP) is rapidly becoming the standard interface for Large Language Models (LLMs) to discover and invoke external t

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

AdaTooler-V: Adaptive Tool-Use for Images and Videos

DGX agent

arXiv:2512.16918v3 Announce Type: replace Abstract: Recent advances have shown that multimodal large language models (MLLMs) benefit from multimodal interleaved chain-of-thought (CoT) with vision tool

model-releasesarxiv-cs-cv
29 Apr 2026
Agents

AgenticQwen: Training Small Agentic Language Models with Dual Data Flywheels for Industrial-Scale Tool Use

DGX agent

arXiv:2604.21590v1 Announce Type: new Abstract: Modern industrial applications increasingly demand language models that act as agents, capable of multi-step reasoning and tool use in real-world settin

agentsarxiv-cs-cl
24 Apr 2026
Safety

Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors

DGX agent

arXiv:2601.15625v2 Announce Type: replace Abstract: Large language models (LLMs) can call tools effectively, yet they remain brittle in multi-turn execution: after a tool-call error, smaller models of

safetyarxiv-cs-lg
21 Apr 2026
Model Releases

Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions

DGX agent

arXiv:2509.18847v3 Announce Type: replace-cross Abstract: Tool-augmented large language models (LLMs) are usually trained with supervised imitation or coarse-grained reinforcement learning that optimi

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks

DGX agent

arXiv:2604.10015v1 Announce Type: new Abstract: Recent studies demonstrate that tool-calling capability enables large language models (LLMs) to interact with external environments for long-horizon fin

model-releasesarxiv-cs-ai
14 Apr 2026
Tutorials

Use of AI Tools: Guidelines to Maintain Academic Integrity in Computing Colleges

DGX agent

arXiv:2604.11111v1 Announce Type: cross Abstract: The rapid adoption of AI tools such as ChatGPT has significantly transformed academic practices, offering considerable benefits for both students and

tutorialsarxiv-cs-ai
14 Apr 2026
Agents

AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP

DGX agent

arXiv:2607.11098v2 Announce Type: replace-cross Abstract: Tool-using LLM agents are mostly evaluated assuming all tools work. When a tool times out, returns a week-stale value, or has its description

agentsarxiv-cs-ai
15 Jul 2026
Agents

LAVE: Latent Visual Evidence-Enhanced Planning for Video Tool-use Agents

DGX agent

arXiv:2608.07585v1 Announce Type: new Abstract: Long-video understanding requires models to efficiently acquire and reuse sparse visual evidence from long and redundant video streams. Recent video too

agentsarxiv-cs-cv
11 Aug 2026
Model Releases

VTO: Visual Tool Orchestration for Video Anomaly Detection

DGX agent

arXiv:2608.08219v1 Announce Type: cross Abstract: Video anomaly detection (VAD) is a critical yet challenging task due to the complex and diverse nature of real-world scenarios. Traditional deep learn

model-releasesarxiv-cs-ai
11 Aug 2026
Applications

Taxonomy-Driven Analysis of Open-Source AI Risk Mitigation Tools

DGX agent

arXiv:2608.07446v1 Announce Type: cross Abstract: Rapid adoption of large language models (LLMs) in enterprise settings has introduced operational, security, and governance risks. As generative AI app

applicationsarxiv-cs-ai
10 Aug 2026
Model Releases

UrbanAgent: A Tool-Augmented Agent for Cross-System Urban Tasks

DGX agent

arXiv:2608.03018v1 Announce Type: new Abstract: Modern cities rely on an increasing number of digital services to operate, but residents' daily needs are still difficult to meet. Services are fragment

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios

DGX agent

arXiv:2607.23722v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as agents that interact with stateful environments over multiple steps: gathering hidden informat

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

How to Realize Recursively Self-Improving Agents and Personal Singularity: A Goal-, Scope-, Tool-, and Benchmark-Driven Multi-Agent Architecture

DGX agent

arXiv:2607.12254v1 Announce Type: new Abstract: Large language model (LLM) agents can increasingly plan, use tools, maintain memory, and execute long-horizon tasks. These advances motivate two linked

model-releasesarxiv-cs-cv
15 Jul 2026
Agents

Tool-Making and Self-Evolving LLM Agents in Low-Latency Systems

DGX agent

arXiv:2607.08010v1 Announce Type: new Abstract: Production LLM agents often waste latency and reliability by regenerating code for the same procedural steps on every request. We replace this inference

agentsarxiv-cs-cl
10 Jul 2026
Model Releases

Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents

DGX agent

arXiv:2607.07405v1 Announce Type: new Abstract: Tool-using LLM agents can violate the very policies they are deployed to enforce while appearing to complete the task successfully. In policy-permissive

model-releasesarxiv-cs-ai
9 Jul 2026
Agents

When Does Tool Use Increase the Expressive Power of Finite-Precision Recurrent Models?

DGX agent

arXiv:2607.06155v1 Announce Type: cross Abstract: Modern sequence models are increasingly deployed as agents that interleave token generation with calls to external tools. We give an exact, architectu

agentsarxiv-cs-cl
8 Jul 2026
Model Releases

When Users Are Happy but Agents Are Wrong: Multi-Dimensional Evaluation of Tool-Augmented Dialogue

DGX agent

arXiv:2510.19186v3 Announce Type: replace Abstract: Evaluating conversational AI systems that use external tools is challenging, as errors can arise from complex interactions among user, agent, and to

model-releasesarxiv-cs-cl
7 Jul 2026
Model Releases

PHREEQC-MCQ-200: A Diagnostic Benchmark for Tool-Augmented Scientific Simulator Agents

DGX agent

arXiv:2607.00436v1 Announce Type: new Abstract: Large language model agents are increasingly connected to scientific software, yet it remains unclear when tool access makes scientific computation more

model-releasesarxiv-cs-ai
2 Jul 2026
Safety

Agentic Tool Use in Large Language Models

DGX agent

arXiv:2604.00835v2 Announce Type: replace Abstract: Large language models are increasingly being deployed as autonomous agents yet their real world effectiveness depends on reliable tools for informat

safetyarxiv-cs-cl
30 Jun 2026
Model Releases

Dynamo: Dynamic Skill-Tool Evolution for Vision-Language Agents

DGX agent

arXiv:2606.30185v1 Announce Type: new Abstract: Improving vision-language models (VLMs) on visual reasoning typically requires retraining or hand-designed prompts and tools. We present Dynamo, a train

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Less Context, Better Agents: Efficient Context Engineering for Long-Horizon Tool-Using LLM Agents

DGX agent

arXiv:2606.10209v1 Announce Type: new Abstract: Large language models deployed as autonomous agents for enterprise workflows face a key challenge: verbose tool responses from enterprise systems can ca

model-releasesarxiv-cs-ai
10 Jun 2026
Agents

AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning

DGX agent

arXiv:2512.13278v2 Announce Type: replace Abstract: Agentic reinforcement learning has advanced large language models (LLMs) to reason through long chain-of-thought trajectories while interleaving ext

agentsarxiv-cs-cl
8 Jun 2026
Safety

DexFuture: Hierarchical Future-State Visuomotor Targeting for Bimanual Dexterous Tool Use

DGX agent

arXiv:2606.05699v1 Announce Type: new Abstract: Bimanual dexterous tool use remains challenging for robots due to high-dimensional hand configurations and complex hand-tool-object dynamics and contact

safetyarxiv-cs-ro
5 Jun 2026
Model Releases

VESTA: Visual Exploration with Statistical Tool Agents

DGX agent

arXiv:2606.00384v1 Announce Type: new Abstract: Fitting quantitative models to data is a central step in scientific workflows, yet it remains one of the least automated. Recent agent-based systems lev

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Depth-Dependent Indirect Prompt Injection in Tool-Calling ReAct Agents: Injection Depth, Payload Framing, and Turn-Budget Sensitivity

DGX agent

arXiv:2605.30686v1 Announce Type: cross Abstract: ReAct agents that interleave chain-of-thought reasoning with tool calls are increasingly deployed for real tasks such as scheduling, file retrieval, a

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Prompts Don't Protect: Architectural Enforcement via MCP Proxy for LLM Tool Access Control

DGX agent

arXiv:2605.18414v1 Announce Type: cross Abstract: Large language models increasingly operate as autonomous agents that select and invoke tools from large registries. We identify a critical gap: when u

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents

DGX agent

arXiv:2605.15104v1 Announce Type: new Abstract: Voice agents increasingly require reliable tool use from speech, whereas prominent tool-calling benchmarks remain text-based. We study whether verified

model-releasesarxiv-cs-cl
15 May 2026
Model Releases

Latency-Quality Routing for Functionally Equivalent Tools in LLM Agents

DGX agent

arXiv:2605.14241v1 Announce Type: new Abstract: Tool-augmented LLM agents increasingly access the same tool type through multiple functionally equivalent providers, such as web-search APIs, retrievers

model-releasesarxiv-cs-lg
15 May 2026
Agents

MCPShield: Content-Aware Attack Detection for LLM Agent Tool-Call Traffic

DGX agent

arXiv:2605.11053v1 Announce Type: cross Abstract: The Model Context Protocol (MCP) has become a widely adopted interface for LLM agents to invoke external tools, yet learned monitoring of MCP tool-cal

agentsarxiv-cs-lg
13 May 2026
Agents

Agent-First Tool API: A Semantic Interface Paradigm for Enterprise AI Agent Systems

DGX agent

arXiv:2605.10555v1 Announce Type: new Abstract: As AI agents transition from research prototypes to enterprise production systems, the tool interfaces they consume remain rooted in human-oriented CRUD

agentsarxiv-cs-ai
12 May 2026
Agents

Simulating Complex Multi-Turn Tool Calling Interactions in Stateless Execution Environments

DGX agent

arXiv:2601.19914v2 Announce Type: replace-cross Abstract: Synthetic data has proven itself to be a valuable resource for tuning smaller, cost-effective language models to handle the complexities of mu

agentsarxiv-cs-ai
12 May 2026
Model Releases

TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning

DGX agent

arXiv:2605.09544v1 Announce Type: new Abstract: Tool-integrated reasoning has emerged as a promising paradigm for enhancing large language models with external computation, retrieval, and execution ca

model-releasesarxiv-cs-ai
12 May 2026
← Previous
123456…108
Next →