AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
5,141 results
Model Releases

The Remarkable Effectiveness of Providing AI Agents with Natural Language Tools: A Replication Study Validating NLT Performance Across 14 Models

DGX agent

arXiv:2607.03953v1 Announce Type: cross Abstract: This study independently replicates and extends the Natural Language Tools (NLT) framework of Johnson et al.~(2025), which questions the use of struct

model-releasesarxiv-cs-ai
7 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Local Ai

Looking Is Not Picking: An Attention-Segment Account of Tool-Selection Failures in LLM Agents

DGX agent

arXiv:2606.16364v2 Announce Type: replace Abstract: LLM agents mis-call tools, and the natural guess is that the model failed to see the right tool in a crowded harness. We show the opposite through a

local-aiarxiv-cs-ai
30 Jun 2026
Model Releases

ShareLock: A Stealthy Multi-Tool Threshold Poisoning Attack Against MCP

DGX agent

arXiv:2606.27027v1 Announce Type: cross Abstract: With the rapid evolution of LLM-driven agents, Model Context Protocol (MCP), an open protocol bridging LLMs with external tools, has quickly become fo

model-releasesarxiv-cs-ai
26 Jun 2026
Agents

ToolGate: Token-Efficient Pre-Call Control for Tool-Augmented Vision-Language Agents

DGX agent

arXiv:2606.03054v1 Announce Type: new Abstract: Tool-augmented vision-language agents can acquire external perceptual evidence through OCR, detection, segmentation, and other tools, but executing ever

agentsarxiv-cs-ai
3 Jun 2026
Model Releases

Self-Healing Agentic Orchestrators for Reliable Tool-Augmented Large Language Model Systems

DGX agent

arXiv:2606.01416v1 Announce Type: new Abstract: Tool-augmented large language model (LLM) agents rely on orchestration layers that coordinate planning, retrieval, tool invocation, validation, memory,

model-releasesarxiv-cs-ai
2 Jun 2026
Agents

SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL

DGX agent

arXiv:2512.04069v2 Announce Type: replace Abstract: Vision Language Models (VLMs) demonstrate strong qualitative visual understanding, but struggle with metrically precise spatial reasoning required f

agentsarxiv-cs-cv
2 Jun 2026
Agents

Do Agents Know What They Can't Do? Evaluating Feasibility Awareness in Tool-Using Agents

DGX agent

arXiv:2605.28532v1 Announce Type: new Abstract: Tool-using agents often incur substantial computational cost due to long reasoning chains and iterative tool usage. In practical scenarios, many tasks b

agentsarxiv-cs-ai
28 May 2026
Model Releases

Tool Forge: A Validation-Carrying Toolchain for Governed Agentic Execution

DGX agent

arXiv:2605.28000v1 Announce Type: cross Abstract: Large language model agents are increasingly expected to perform operational work: calling APIs, manipulating files, assembling workflows, and acting

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Memory-Induced Tool-Drift in LLM Agents

DGX agent

arXiv:2605.24941v1 Announce Type: cross Abstract: Modern LLM agents combine long-term memory for personalization with tool-calling interfaces for taking actions in the world -- a combination underpinn

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

ToolWeave: Structured Synthesis of Complex Multi-Turn Tool-Calling Dialogues

DGX agent

arXiv:2605.12521v1 Announce Type: cross Abstract: Multi-turn tool calling is essential for LLMs to function as autonomous agents, yet synthesizing the training data required for these capabilities rem

model-releasesarxiv-cs-ai
14 May 2026
Model Releases

Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models

DGX agent

arXiv:2605.13119v1 Announce Type: cross Abstract: Vision-language-action (VLA) models are effective robot action executors, but they remain limited on long-horizon tasks due to the dual burden of exte

model-releasesarxiv-cs-ai
14 May 2026
Agents

ChemAmp: Amplified Chemistry Tools via Composable Agents

DGX agent

arXiv:2505.21569v3 Announce Type: replace-cross Abstract: Although LLM-based agents are proven to master tool orchestration in scientific fields, particularly chemistry, their single-task performance

agentsarxiv-cs-ai
20 Apr 2026
Agents

ToolSpec: Accelerating Tool Calling via Schema-Aware and Retrieval-Augmented Speculative Decoding

DGX agent

arXiv:2604.13519v1 Announce Type: new Abstract: Tool calling has greatly expanded the practical utility of large language models (LLMs) by enabling them to interact with external applications. As LLM

agentsarxiv-cs-cl
16 Apr 2026
Model Releases

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs

DGX agent

arXiv:2604.12896v1 Announce Type: new Abstract: Multimodal language models (MLLMs) are increasingly paired with vision tools (e.g., depth, flow, correspondence) to enhance visual reasoning. However, d

model-releasesarxiv-cs-cv
15 Apr 2026
Model Releases

Disambiguation-Centric Finetuning Makes Enterprise Tool-Calling LLMs More Realistic and Less Risky

DGX agent

arXiv:2507.03336v4 Announce Type: replace Abstract: Large language models (LLMs) are increasingly tasked with invoking enterprise APIs, yet they routinely falter when near-duplicate tools vie for the

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

Tool Retrieval Bridge: Aligning Vague Instructions with Retriever Preferences via Bridge Model

DGX agent

arXiv:2604.07816v1 Announce Type: new Abstract: Tool learning has emerged as a promising paradigm for large language models (LLMs) to address real-world challenges. Due to the extensive and irregularl

model-releasesarxiv-cs-cl
10 Apr 2026
Model Releases

Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks

DGX agent

arXiv:2608.10357v1 Announce Type: cross Abstract: Long-horizon tool-using agents must reason over user goals, domain policies, tool calls, simulator state, and delayed verifiable rewards. Reinforcemen

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

When History Lies: Evaluating and Improving Tool Use under Misleading Multi-Turn Histories

DGX agent

arXiv:2608.06057v1 Announce Type: new Abstract: Tool-calling agents infer task state from accumulated dialogue and tool traces. In persistent interactions, however, historical traces may remain struct

model-releasesarxiv-cs-ai
7 Aug 2026
Agents

Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance

DGX agent

arXiv:2608.03403v1 Announce Type: new Abstract: The performance bottleneck of agents is increasingly shifting from model capability to the robustness of their execution processes. Tools play a central

agentsarxiv-cs-ai
5 Aug 2026
Agents

Verified Tool Calls Improve LLM Agent Reliability Under Non-Atomic Failures

DGX agent

arXiv:2608.02645v1 Announce Type: cross Abstract: Large Language Model (LLM) agents rely on external tools to perform multistage tasks. Existing agent frameworks typically assume that tool calls are a

agentsarxiv-cs-ai
5 Aug 2026
Model Releases

Determinants and Limits of LLM Security-Tool Orchestration: A Study with HexStrike-AI

DGX agent

arXiv:2607.02873v1 Announce Type: cross Abstract: Large language model agents driving security tool suites over the Model Context Protocol are increasingly common. Yet the factors that bound their cap

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

Governed MCP: Kernel-Level Tool Governance for AI Agents via Logit-Based Safety Primitives

DGX agent

arXiv:2604.16870v2 Announce Type: replace-cross Abstract: AI agents increasingly call external tools (file system, network, APIs) through the Model Context Protocol (MCP). These tool calls are the age

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

Expert Evaluation of Clinical AI Tools on Real Point-of-Care Clinical Queries

DGX agent

arXiv:2606.28960v1 Announce Type: new Abstract: Physicians now pose millions of clinical questions to AI tools each week, yet these tools are evaluated largely on hypothetical or exam-style questions,

model-releasesarxiv-cs-ai
30 Jun 2026
Agents

GROW^2: Grounding Which and Where for Robot Tool Use

DGX agent

arXiv:2606.30632v1 Announce Type: cross Abstract: Can the robot use a plate to cut a cake if no knife is available? Tool use greatly expands robot capabilities, but to use tools creatively beyond thei

agentsarxiv-cs-ai
30 Jun 2026
Research

Geometric Reconstruction of Extrinsic Contact Trajectories using Tactile Sensing and Proprioception for Tool Manipulation

DGX agent

arXiv:2606.22251v1 Announce Type: new Abstract: Tactile sensing enables robots to perceive rich contact information at the grasp, supporting tasks such as object recognition, in-hand pose estimation,

researcharxiv-cs-ro
23 Jun 2026
Model Releases

MedCTA: A Benchmark for Clinical Tool Agents

DGX agent

arXiv:2606.11702v1 Announce Type: cross Abstract: To make clinically grounded decisions, medical AI agents are expected to go beyond simple recognition and be capable of tool retrieval, evidence acqui

model-releasesarxiv-cs-ai
11 Jun 2026
Model Releases

Beyond APIs: Probing the Limits of MLLMs in Physical Tool Use

DGX agent

arXiv:2606.10803v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) excel at utilizing digital APIs and increasingly serve as the 'brain' of embodied AI, instructing robots to i

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

The Impact of Configuring Agentic AI Coding Tools on Build-vs-Buy Decisions: A Study Protocol

DGX agent

arXiv:2606.03907v1 Announce Type: cross Abstract: Agentic AI coding tools write code with increasing autonomy and in doing so decide when to import a library and when to implement functionality from s

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Same Payload, Different Channel: Measuring Trust Asymmetry in Tool-Using Language Models

DGX agent

arXiv:2606.00566v1 Announce Type: cross Abstract: As language models take on agentic roles that span calling external APIs, reading tool outputs, and acting on instructions embedded in third-party con

model-releasesarxiv-cs-cl
2 Jun 2026
Safety

MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning

DGX agent

arXiv:2605.26154v1 Announce Type: cross Abstract: LLM-driven agents are capable of selecting external tools to complete users' tasks. However, attackers could compromise such process, steering agents

safetyarxiv-cs-ai
27 May 2026
Agents

Attested Tool-Server Admission: A Security Extension to the Model Context Protocol

DGX agent

arXiv:2605.24248v1 Announce Type: cross Abstract: The Model Context Protocol (MCP) standardizes how a large-language-model (LLM) agent and an external tool server exchange messages, but not trust: a h

agentsarxiv-cs-ai
26 May 2026
Model Releases

TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents

DGX agent

arXiv:2605.16909v1 Announce Type: new Abstract: Tool-using agents are increasingly expected to operate across realistic professional workflows, where they must interpret multimodal inputs, coordinate

model-releasesarxiv-cs-ai
19 May 2026
Safety

Quantitative Certification of Agentic Tool Selection

DGX agent

arXiv:2510.03992v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed in agentic systems, where a fundamental task is mapping user intents to relevant extern

safetyarxiv-cs-ai
14 May 2026
Model Releases

Beyond the Black Box: Interpretability of Agentic AI Tool Use

DGX agent

arXiv:2605.06890v1 Announce Type: new Abstract: AI agents are promising for high-stakes enterprise workflows, but dependable deployment remains limited because tool-use failures are difficult to diagn

model-releasesarxiv-cs-ai
11 May 2026
Agents

Position: Agent Should Invoke External Tools ONLY When Epistemically Necessary

DGX agent

arXiv:2506.00886v3 Announce Type: replace Abstract: As large language models evolve into tool-augmented agents, a central question remains unresolved: when is external tool use actually justified? Exi

agentsarxiv-cs-ai
7 May 2026
Model Releases

TSCG: Deterministic Tool-Schema Compilation for Agentic LLM Deployments

DGX agent

arXiv:2605.04107v1 Announce Type: cross Abstract: Production agent frameworks (OpenAI Function Calling, Anthropic Tool Use, MCP) transmit tool schemas as JSON, a format designed for machine parsing, n

model-releasesarxiv-cs-cl
7 May 2026
Agents

SCRIBE: Structured Mid-Level Supervision for Tool-Using Language Models

DGX agent

arXiv:2601.03555v2 Announce Type: replace Abstract: Training reliable tool-augmented agents remains a significant challenge, largely due to the difficulty of credit assignment in multi-step reasoning.

agentsarxiv-cs-ai
28 Apr 2026
Model Releases

Trajectory2Task: Training Robust Tool-Calling Agents with Synthesized Yet Verifiable Data for Complex User Intents

DGX agent

arXiv:2601.20144v3 Announce Type: replace Abstract: Tool-calling agents are increasingly deployed in real-world customer-facing workflows. Yet most studies on tool-calling agents focus on idealized se

model-releasesarxiv-cs-cl
23 Apr 2026
Model Releases

Latent Preference Modeling for Cross-Session Personalized Tool Calling

DGX agent

arXiv:2604.17886v1 Announce Type: new Abstract: Users often omit essential details in their requests to LLM-based agents, resulting in under-specified inputs for tool use. This poses a fundamental cha

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

The Reasoning Trap: How Enhancing LLM Reasoning Amplifies Tool Hallucination

DGX agent

arXiv:2510.22977v2 Announce Type: replace-cross Abstract: Enhancing the reasoning capabilities of Large Language Models (LLMs) is a key strategy for building Agents that 'think then act.' However, rec

model-releasesarxiv-cs-ai
20 Apr 2026
Local Ai

ClawGuard: A Runtime Security Framework for Tool-Augmented LLM Agents Against Indirect Prompt Injection

DGX agent

arXiv:2604.11790v1 Announce Type: cross Abstract: Tool-augmented Large Language Model (LLM) agents have demonstrated impressive capabilities in automating complex, multi-step real-world tasks, yet rem

local-aiarxiv-cs-ai
14 Apr 2026
Model Releases

Benchmarking LLM Tool-Use in the Wild

DGX agent

arXiv:2604.06185v1 Announce Type: cross Abstract: Fulfilling user needs through Large Language Model multi-turn, multi-step tool-use is rarely a straightforward process. Real user interactions are inh

model-releasesarxiv-cs-ai
10 Apr 2026
Safety

Learning to Coordinate Symbolic Tools: LLM Agents for Verified Sum-of-Squares Certificates

DGX agent

arXiv:2608.00326v2 Announce Type: replace Abstract: Tool calling allows large language models (LLMs) to invoke external computation during problem solving, a useful capability in various fields includ

safetyarxiv-cs-ai
11 Aug 2026
Model Releases

Models as Tools: An Agentic Coordination Framework for Unified Multimodal Visual Tracking

DGX agent

arXiv:2608.00847v1 Announce Type: new Abstract: Most current visual trackers adopt a matching-based architecture trained exclusively on tracking datasets, whose performance gains depend heavily on the

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

MagicSelector: Joint Optimization for Agent Tool Selection via Counterfactual Decomposition and Progressive Reranking

DGX agent

arXiv:2607.17751v2 Announce Type: cross Abstract: We present MagicSelector, a joint optimization framework integrating Counterfactual task decomposition, Progressive reranking, and Dynamic Top-K, desi

model-releasesarxiv-cs-cl
31 Jul 2026
Safety

Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents

DGX agent

arXiv:2607.19449v1 Announce Type: cross Abstract: Evaluation frameworks for tool-augmented LLM agents focus overwhelmingly on capability metrics or explicit tool crashes, leaving silent infrastructure

safetyarxiv-cs-ai
23 Jul 2026
Model Releases

Context-Aware Force Estimation for Deformable Tool Manipulation in Robotic Environmental Swabbing via Few-Shot Continual Adaptation

DGX agent

arXiv:2607.07574v1 Announce Type: new Abstract: Robotic surface swabbing requires sustained interaction between a compliant tool and heterogeneous environments, where accurate estimation of tip-level

model-releasesarxiv-cs-ro
9 Jul 2026
Safety

PORTS: Preference-Optimized Retrievers for Tool Selection with Large Language Models

DGX agent

arXiv:2607.05441v1 Announce Type: cross Abstract: Integrating external tools with Large Language Models (LLMs) has emerged as a promising paradigm for accomplishing complex tasks. Since LLMs still str

safetyarxiv-cs-ai
8 Jul 2026
← Previous
12345…108
Next →