AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
61+ results
5 Aug 2026

HyperAgent: Planning and Acting over Tool-Schema Hypergraphs for Tool-Use LLM Agents

AgentsDGX agent

arXiv:2608.02650v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on external tools to complete complex real-world tasks. However, reliable tool-use planning remains

ToolLIFT: Lifting Tool-Specific Trajectories into Function-Level Graphs for Generalizable Tool Planning

AgentsDGX agent

arXiv:2608.03468v1 Announce Type: new Abstract: Historical tool-use trajectories provide valuable experience for large language model (LLM) agents to plan and coordinate tool usage. Existing approache

Screenshots or Tools? Eliciting Tool Use and Managing Multimodal Context in Hybrid GUI-MCP Computer-Use Agents

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.03327v1 Announce Type: new Abstract: Hybrid computer-use agents can act through screenshots or call text tools. We find that having a tool available does not settle which way the effect goe

30 Jul 2026

Under the Hood: Serving Kimi K3

Model ReleasesDGX agent

DigitalOcean launched Kimi K3 on day 0. It’s already one of the most popular models on the platform and across the market: second most likes on Hugging Face, sixth most traffic on OpenCode. Getting a

11 May 2026

GTIG AI Threat Tracker: Adversaries Leverage AI for Vulnerability Exploitation, Augmented Operations, and Initial Access

Model ReleasesDGX agent

Executive Summary Since our February 2026 report on AI-related threat activity, Google Threat Intelligence Group (GTIG) has continued to track a maturing transition from nascent AI-enabled operations

ARMOR: An Agentic Framework for Reaction Feasibility Prediction via Adaptive Utility-aware Multi-tool Reasoning

AgentsDGX agent

arXiv:2605.07103v1 Announce Type: new Abstract: Reaction feasibility prediction, as a fundamental problem in computational chemistry, has benefited from diverse tools enabled by recent advances in art

19 May 2026

ToolMATH: A Diagnostic Benchmark for Long-Horizon Tool Use under Systematic Tool-Catalog Constraints

Model ReleasesDGX agent

arXiv:2602.21265v2 Announce Type: replace Abstract: We introduce ToolMATH, a math-grounded diagnostic benchmark for evaluating long-horizon tool use under controllable tool-catalog conditions. ToolMAT

Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback

Model ReleasesDGX agent

arXiv:2605.17453v1 Announce Type: cross Abstract: Tool-using LLM agents increasingly rely on external tools to make consequential decisions, yet most existing agent-security benchmarks and defenses im

Implicit Hierarchical GRPO: Decoupling Tool Invocation from Execution for Tool-Integrated Mathematical Reasoning

SafetyDGX agent

arXiv:2605.18500v1 Announce Type: new Abstract: Large language models (LLMs) have increasingly leveraged tool invocation to enhance their reasoning capabilities. However, existing approaches typically

21 Jul 2026

A Fireside Chat with Cat and Thariq from the Claude Code team

Model ReleasesDGX agent

Earlier this month I hosted a fireside chat session at the AI Engineer World's Fair with Cat Wu and Thariq Shihipar from Anthropic's Claude Code team. We talked about Claude Code, Claude Tag, Fable, c

10 Apr 2026

When to Trust Tools? Adaptive Tool Trust Calibration For Tool-Integrated Math Reasoning

Model ReleasesDGX agent

arXiv:2604.08281v1 Announce Type: new Abstract: Large reasoning models (LRMs) have achieved strong performance enhancement through scaling test time computation, but due to the inherent limitations of

23 Apr 2026

The Tool-Overuse Illusion: Why Does LLM Prefer External Tools over Internal Knowledge?

SafetyDGX agent

arXiv:2604.19749v1 Announce Type: new Abstract: Equipping LLMs with external tools effectively addresses internal reasoning limitations. However, it introduces a critical yet under-explored phenomenon

JTPRO: A Joint Tool-Prompt Reflective Optimization Framework for Language Agents

AgentsDGX agent

arXiv:2604.19821v1 Announce Type: new Abstract: Large language model (LLM) agents augmented with external tools often struggle as number of tools grow large and become domain-specific. In such setting

12 Aug 2026

DOCSCHISEL: Adaptive Tool Documentation Optimization Framework for LLM Agents

AgentsDGX agent

arXiv:2608.10037v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly rely on external tools to accomplish complex real-world tasks, making tool documentation a critical groundin

12 May 2026

PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning

TutorialsDGX agent

arXiv:2605.09931v1 Announce Type: cross Abstract: Tool-integrated reasoning (TIR) enables large language models (LLMs) to enhance their capabilities by interacting with external tools, such as code in

ToolScope: Enhancing LLM Agent Tool Use through Tool Merging and Context-Aware Filtering

AgentsDGX agent

arXiv:2510.20036v2 Announce Type: replace Abstract: Large language model (LLM) agents rely on external tools to solve complex tasks, but real-world toolsets often contain redundant tools with overlapp

TRACER: Verifiable Generative Provenance for Multimodal Tool-Using Agents

Model ReleasesDGX agent

arXiv:2605.09934v1 Announce Type: new Abstract: Multimodal large language models increasingly solve vision-centric tasks by calling external tools for visual inspection, OCR, retrieval, calculation, a

27 May 2026

Mind the Tool Failures: Achieving Synergistic Tool Gains for Medical Agents

AgentsDGX agent

arXiv:2605.26691v1 Announce Type: new Abstract: Medical AI agents increasingly use external tools for diagnosis, treatment recommendation, and evidence retrieval, yet most existing approaches assume t

21 Apr 2026

Don't Adapt Small Language Models for Tools; Adapt Tool Schemas to the Models

AgentsDGX agent

arXiv:2510.07248v3 Announce Type: replace Abstract: Small language models (SLMs) enable scalable tool-augmented multi-agent systems where multiple SLMs handle subtasks orchestrated by a powerful coord

29 Jul 2026

Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction

AgentsDGX agent

arXiv:2607.25718v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly rely on invoking external tools to complete real-world tasks. Tool retrieval, which selects a small tas

9 Jun 2026

Initial impressions of Claude Fable 5

Model ReleasesDGX agent

I didn't have early access to today's Claude Fable 5 release, but I've spent the past ~5.5 hours putting it through its paces. My initial impressions are that this is something of a beast. It's slow,

Contract2Tool: Learning Preconditions and Effects for Reliable Tool-Augmented LLM Agents

AgentsDGX agent

arXiv:2606.07904v1 Announce Type: new Abstract: Tool-augmented large language model agents increasingly rely on external APIs, but standard tool schemas describe how to call a tool, not when the tool

15 May 2026

Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use

AgentsDGX agent

arXiv:2605.14038v1 Announce Type: new Abstract: Large language models (LLMs) increasingly act as autonomous agents that must decide when to answer directly vs. when to invoke external tools. Prior wor

14 May 2026

ReTool-Video: Recursive Tool-Using Video Agents with Meta-Augmented Tool Grounding

Model ReleasesDGX agent

arXiv:2605.13228v1 Announce Type: cross Abstract: Video understanding requires active evidence seeking, motivating tool-augmented video agents for temporal reasoning, cross-modal understanding, and co

17 Apr 2026

RaTA-Tool: Retrieval-based Tool Selection with Multimodal Large Language Models

SafetyDGX agent

arXiv:2604.14951v1 Announce Type: cross Abstract: Tool learning with foundation models aims to endow AI systems with the ability to invoke external resources -- such as APIs, computational utilities,

16 Apr 2026

GLM-5.1 Tool Calling Issue Fix & Chat Template Update If you are running GLM-5.1 with vLLM/SGLang and using tool calling, please update your…

Model ReleasesDGX agent

GLM-5.1 Tool Calling Issue Fix & Chat Template Update If you are running GLM-5.1 with vLLM/SGLang and using tool calling, please update your chat template. http://huggingface.co/zai-org/GLM-5.1/blob/m

ToolOmni: Enabling Open-World Tool Use via Agentic learning with Proactive Retrieval and Grounded Execution

Model ReleasesDGX agent

arXiv:2604.13787v1 Announce Type: new Abstract: Large Language Models (LLMs) enhance their problem-solving capability by utilizing external tools. However, in open-world scenarios with massive and evo

8 Apr 2026

Meta's new model is Muse Spark, and meta.ai chat has some interesting tools

Model ReleasesDGX agent

Meta announced Muse Spark today, their first model release since Llama 4 almost exactly a year ago. It's hosted, not open weights, and the API is currently 'a private API preview to select users', but

25 Jun 2026

Beyond Function Calling: Benchmarking Tool-Using Agents under Tool-Environment Unreliability

Model ReleasesDGX agent

arXiv:2606.25819v1 Announce Type: new Abstract: Large language models are increasingly deployed as agents that solve tasks by interacting with external tool environments. Although recent tool-use benc

21 May 2026

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning

AgentsDGX agent

arXiv:2605.20342v1 Announce Type: new Abstract: Training large multimodal models (LMMs) via reinforcement learning (RL) to natively invoke video-processing tools (e.g., cropping) has become a promisin

15 Apr 2026

Track usage and costs across your users, agents and tools. We're also shipping better access controls for your tools & agents: - Set spend l…

AgentsDGX agent

Track usage and costs across your users, agents and tools. We're also shipping better access controls for your tools & agents: - Set spend limits on agents - Set spend limits for users - Control what

New in LangSmith Fleet: Tool access controls and usage tracking. 📊 Track cost and usage by user, agent, and tool from a single dashboard 💳…

AgentsDGX agent

New in LangSmith Fleet: Tool access controls and usage tracking. 📊 Track cost and usage by user, agent, and tool from a single dashboard 💳 Set spend limits per user or team to prevent surprises ✅ Cont

14 Apr 2026

Do LLMs Know Tool Irrelevance? Demystifying Structural Alignment Bias in Tool Invocations

SafetyDGX agent

arXiv:2604.11322v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated impressive capabilities in utilizing external tools. In practice, however, LLMs are often exposed to to

TInR: Exploring Tool-Internalized Reasoning in Large Language Models

SafetyDGX agent

arXiv:2604.10788v1 Announce Type: cross Abstract: Tool-Integrated Reasoning (TIR) has emerged as a promising direction by extending Large Language Models' (LLMs) capabilities with external tools durin

☁️ Salesforce tools now in Fleet One of the most requested features we've gotten, and it's now a first-class supported tool in Fleet! Just s…

AgentsDGX agent

☁️ Salesforce tools now in Fleet One of the most requested features we've gotten, and it's now a first-class supported tool in Fleet! Just sign in with your Salesforce account, and start using it imme

10 Aug 2026

// The Bitter Lesson of Tool Calling // Tool calling is a design choice, and the defaults are quietly costing accuracy. How so? New research…

Model ReleasesDGX agent

// The Bitter Lesson of Tool Calling // Tool calling is a design choice, and the defaults are quietly costing accuracy. How so? New research releases a generation-spanning comparison of programmatic t

6 Aug 2026

Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools

Model ReleasesDGX agent

arXiv:2608.04719v1 Announce Type: new Abstract: Agent evaluations tell us that a model picked the wrong tool, but rarely why. We introduce canary tools: diagnostic probe tools planted in an agent's Mo

22 May 2026

What Does Vision Tool-Use Reinforcement Learning Really Learn? Disentangling Tool-Induced and Intrinsic Effects for Crop-and-Zoom

AgentsDGX agent

arXiv:2602.01334v2 Announce Type: replace Abstract: Vision tool-use reinforcement learning (RL) can equip vision language models with visual operators such as crop-and-zoom and achieves strong perform

30 Apr 2026

Learning to Rewrite Tool Descriptions for Reliable LLM-Agent Tool Use

AgentsDGX agent

arXiv:2602.20426v2 Announce Type: replace Abstract: While most efforts to improve LLM-based tool-using agents focus on the agent itself - through larger models, better prompting, or fine-tuning - agen

11 Aug 2026

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning

AgentsDGX agent

arXiv:2608.09682v1 Announce Type: new Abstract: Tool-augmented vision-language models increasingly 'think with images': they call crop, zoom, or code tools and reason over the returned pixels. However

4 Aug 2026

A Few Neurons Reveal When LLMs Misuse Tools: Sparse Detection and Selective Steering for Reliable Tool Use

Model ReleasesDGX agent

arXiv:2608.00218v1 Announce Type: new Abstract: Agentic LLMs exhibit three consequential tool-use failures: invalid arguments (validity), unnecessary calls (over-calling), and omitted calls when tools

New release of LLM adds support for reasoning traces, OpenAI Responses, server-side tools, and smarter logging

Model ReleasesDGX agent

I released LLM 0.32 this morning, the most significant new version of LLM since the initial launch of the project. The new version includes support for visible reasoning traces, server-side provider t

Control Under Compression: Reliability Frontiers for Tool-Using Agents

Model ReleasesDGX agent

arXiv:2608.01056v1 Announce Type: cross Abstract: Tool-using language-model agents are governed not only by task prompts but also by persistent system-side instructions that specify tools, arguments,

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise

Model ReleasesDGX agent

arXiv:2608.02372v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in task-oriented dialogue systems that support multi-step decision-making in high-stakes domains

6 Jun 2026

ToolChoiceConfusion: Causal Minimal Tool Filtering for Reliable LLM Agents

Model ReleasesDGX agent

arXiv:2606.06284v1 Announce Type: new Abstract: Large language model agents increasingly rely on external tools, but larger tool menus can reduce reliability and efficiency by increasing wrong-tool ca

7 Jul 2026

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents

Model ReleasesDGX agent

arXiv:2607.04686v1 Announce Type: cross Abstract: Tool calling is central to modern language model agents, but aggregate benchmark scores often hide where tool use fails. A model that never calls a ne

26 May 2026

How Many Tools Should an LLM Agent See? A Chance-Corrected Answer

Model ReleasesDGX agent

arXiv:2605.24660v1 Announce Type: cross Abstract: Before an LLM agent can use a tool, a retrieval system must decide which candidate tools to show to the agent. How long should that shortlist be? Show

24 Apr 2026

Tool Attention Is All You Need: Dynamic Tool Gating and Lazy Schema Loading for Eliminating the MCP/Tools Tax in Scalable Agentic Workflows

Model ReleasesDGX agent

arXiv:2604.21816v1 Announce Type: new Abstract: The Model Context Protocol (MCP) has become a common interface for connecting large language model (LLM) agents to external tools, but its reliance on s

2 Jun 2026

Do Multimodal Agents Really Benefit from Tool Use? A Systematic Study of Capability Gains

Model ReleasesDGX agent

arXiv:2606.02357v1 Announce Type: cross Abstract: Tool-augmented multimodal agents show strong benchmark gains, often taken as evidence that agents have learned to use tools. We argue that this interp

Ghost Tool Calls: Issue-Time Privacy for Speculative Agent Tools

AgentsDGX agent

arXiv:2606.02483v1 Announce Type: cross Abstract: Tool-augmented language agents speculatively issue likely future tool calls to hide latency, but those calls leak inferred user intent to external ser

5 May 2026

Are Tools All We Need? Unveiling the Tool-Use Tax in LLM Agents

AgentsDGX agent

arXiv:2605.00136v1 Announce Type: new Abstract: Tool-augmented reasoning has become a popular direction for LLM-based agents, and it is widely assumed to improve reasoning and reliability. However, we

31 Jul 2026

FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification

AgentsDGX agent

arXiv:2607.28225v1 Announce Type: new Abstract: Agentic vision-language models (VLMs), which interleave textual reasoning with explicit tool calls such as cropping and code-based image manipulation, h

20 May 2026

Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning

Model ReleasesDGX agent

arXiv:2605.19852v1 Announce Type: new Abstract: Tool-augmented reasoning has emerged as a promising direction for enhancing the reasoning capabilities of multimodal large language models (MLLMs). Howe

26 Apr 2026

Opened a llama.cpp discussion about whether custom GBNF grammars can compose with tool calls in llama-server. Right now tools work alone, gr…

Model ReleasesDGX agent

Opened a llama.cpp discussion about whether custom GBNF grammars can compose with tool calls in llama-server. Right now tools work alone, grammar works alone, but tools+grammar doesn't. If you use lla

10 Jul 2026

A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding

Local AiDGX agent

arXiv:2512.21414v2 Announce Type: replace Abstract: Recent tool-use frameworks powered by vision-language models (VLMs) improve image understanding by grounding model predictions with specialized tool

22 Apr 2026

From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models

Model ReleasesDGX agent

arXiv:2511.10899v2 Announce Type: replace Abstract: Tool-augmented Language Models (TaLMs) can invoke external tools to solve problems beyond their parametric capacity. However, it remains unclear whe

15 Jul 2026

Amplitude-Only FFN Intervention for Tool-Structured LLM Inference Method: Gated Evaluation Protocol, and Cross-Model Empirical Results

AgentsDGX agent

arXiv:2607.11183v2 Announce Type: replace Abstract: Large language models increasingly operate as tool-using agents, where small format, argument, or function-call errors can invalidate otherwise plau

24 Jun 2026

The US FDA drops an enforcement complaint against Whoop over its blood pressure tracking tool, reversing a July 2025 warning letter; Whoop is updating the tool (Samantha Kelly/Bloomberg)

IndustryDGX agent

Samantha Kelly / Bloomberg: The US FDA drops an enforcement complaint against Whoop over its blood pressure tracking tool, reversing a July 2025 warning letter; Whoop is updating the tool — The US Foo

18 May 2026

Beyond the Query: 5 Scenarios Laying the Foundation for the Agentic Era

Model ReleasesDGX agent

Accessing enterprise data is shifting from static reports to dynamic use by autonomous systems. To keep up, organizations must route fragmented data from SaaS, IoT, and legacy sources into secure, sca

7 Aug 2026

The Bitter Lesson of Tool Calling

Model ReleasesDGX agent

arXiv:2608.06370v1 Announce Type: new Abstract: Tool use transforms LLMs into agents that act beyond their training data, and for code-capable models, programmatic tool calling extends this further by

← Previous
123…166
Next →
9,953 results