AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
49+ results
Agents

HyperAgent: Planning and Acting over Tool-Schema Hypergraphs for Tool-Use LLM Agents

DGX agent

arXiv:2608.02650v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on external tools to complete complex real-world tasks. However, reliable tool-use planning remains

agentsarxiv-cs-ai
5 Aug 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Under the Hood: Serving Kimi K3

DGX agent

DigitalOcean launched Kimi K3 on day 0. It’s already one of the most popular models on the platform and across the market: second most likes on Hugging Face, sixth most traffic on OpenCode. Getting a

model-releasesdigitalocean
30 Jul 2026
Model Releases

GTIG AI Threat Tracker: Adversaries Leverage AI for Vulnerability Exploitation, Augmented Operations, and Initial Access

DGX agent

Executive Summary Since our February 2026 report on AI-related threat activity, Google Threat Intelligence Group (GTIG) has continued to track a maturing transition from nascent AI-enabled operations

model-releasesgoogle-cloud-ai
11 May 2026
Model Releases

ToolMATH: A Diagnostic Benchmark for Long-Horizon Tool Use under Systematic Tool-Catalog Constraints

DGX agent

arXiv:2602.21265v2 Announce Type: replace Abstract: We introduce ToolMATH, a math-grounded diagnostic benchmark for evaluating long-horizon tool use under controllable tool-catalog conditions. ToolMAT

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

A Fireside Chat with Cat and Thariq from the Claude Code team

DGX agent

Earlier this month I hosted a fireside chat session at the AI Engineer World's Fair with Cat Wu and Thariq Shihipar from Anthropic's Claude Code team. We talked about Claude Code, Claude Tag, Fable, c

model-releasessimon-willison
21 Jul 2026
Model Releases

When to Trust Tools? Adaptive Tool Trust Calibration For Tool-Integrated Math Reasoning

DGX agent

arXiv:2604.08281v1 Announce Type: new Abstract: Large reasoning models (LRMs) have achieved strong performance enhancement through scaling test time computation, but due to the inherent limitations of

model-releasesarxiv-cs-cl
10 Apr 2026
Safety

The Tool-Overuse Illusion: Why Does LLM Prefer External Tools over Internal Knowledge?

DGX agent

arXiv:2604.19749v1 Announce Type: new Abstract: Equipping LLMs with external tools effectively addresses internal reasoning limitations. However, it introduces a critical yet under-explored phenomenon

safetyarxiv-cs-ai
23 Apr 2026
Agents

DOCSCHISEL: Adaptive Tool Documentation Optimization Framework for LLM Agents

DGX agent

arXiv:2608.10037v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly rely on external tools to accomplish complex real-world tasks, making tool documentation a critical groundin

agentsarxiv-cs-ai
12 Aug 2026
Agents

ToolLIFT: Lifting Tool-Specific Trajectories into Function-Level Graphs for Generalizable Tool Planning

DGX agent

arXiv:2608.03468v1 Announce Type: new Abstract: Historical tool-use trajectories provide valuable experience for large language model (LLM) agents to plan and coordinate tool usage. Existing approache

agentsarxiv-cs-ai
5 Aug 2026
Tutorials

PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning

DGX agent

arXiv:2605.09931v1 Announce Type: cross Abstract: Tool-integrated reasoning (TIR) enables large language models (LLMs) to enhance their capabilities by interacting with external tools, such as code in

tutorialsarxiv-cs-ai
12 May 2026
Agents

Mind the Tool Failures: Achieving Synergistic Tool Gains for Medical Agents

DGX agent

arXiv:2605.26691v1 Announce Type: new Abstract: Medical AI agents increasingly use external tools for diagnosis, treatment recommendation, and evidence retrieval, yet most existing approaches assume t

agentsarxiv-cs-ai
27 May 2026
Agents

Don't Adapt Small Language Models for Tools; Adapt Tool Schemas to the Models

DGX agent

arXiv:2510.07248v3 Announce Type: replace Abstract: Small language models (SLMs) enable scalable tool-augmented multi-agent systems where multiple SLMs handle subtasks orchestrated by a powerful coord

agentsarxiv-cs-cl
21 Apr 2026
Model Releases

Screenshots or Tools? Eliciting Tool Use and Managing Multimodal Context in Hybrid GUI-MCP Computer-Use Agents

DGX agent

arXiv:2608.03327v1 Announce Type: new Abstract: Hybrid computer-use agents can act through screenshots or call text tools. We find that having a tool available does not settle which way the effect goe

model-releasesarxiv-cs-ai
5 Aug 2026
Agents

Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction

DGX agent

arXiv:2607.25718v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly rely on invoking external tools to complete real-world tasks. Tool retrieval, which selects a small tas

agentsarxiv-cs-ai
29 Jul 2026
Model Releases

Initial impressions of Claude Fable 5

DGX agent

I didn't have early access to today's Claude Fable 5 release, but I've spent the past ~5.5 hours putting it through its paces. My initial impressions are that this is something of a beast. It's slow,

model-releasessimon-willison
9 Jun 2026
Agents

Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use

DGX agent

arXiv:2605.14038v1 Announce Type: new Abstract: Large language models (LLMs) increasingly act as autonomous agents that must decide when to answer directly vs. when to invoke external tools. Prior wor

agentsarxiv-cs-ai
15 May 2026
Model Releases

ReTool-Video: Recursive Tool-Using Video Agents with Meta-Augmented Tool Grounding

DGX agent

arXiv:2605.13228v1 Announce Type: cross Abstract: Video understanding requires active evidence seeking, motivating tool-augmented video agents for temporal reasoning, cross-modal understanding, and co

model-releasesarxiv-cs-ai
14 May 2026
Safety

RaTA-Tool: Retrieval-based Tool Selection with Multimodal Large Language Models

DGX agent

arXiv:2604.14951v1 Announce Type: cross Abstract: Tool learning with foundation models aims to endow AI systems with the ability to invoke external resources -- such as APIs, computational utilities,

safetyarxiv-cs-cl
17 Apr 2026
Model Releases

GLM-5.1 Tool Calling Issue Fix & Chat Template Update If you are running GLM-5.1 with vLLM/SGLang and using tool calling, please update your…

DGX agent

GLM-5.1 Tool Calling Issue Fix & Chat Template Update If you are running GLM-5.1 with vLLM/SGLang and using tool calling, please update your chat template. http://huggingface.co/zai-org/GLM-5.1/blob/m

model-releaseszhipu-ai--x
16 Apr 2026
Model Releases

Meta's new model is Muse Spark, and meta.ai chat has some interesting tools

DGX agent

Meta announced Muse Spark today, their first model release since Llama 4 almost exactly a year ago. It's hosted, not open weights, and the API is currently 'a private API preview to select users', but

model-releasessimon-willison
8 Apr 2026
Model Releases

Beyond Function Calling: Benchmarking Tool-Using Agents under Tool-Environment Unreliability

DGX agent

arXiv:2606.25819v1 Announce Type: new Abstract: Large language models are increasingly deployed as agents that solve tasks by interacting with external tool environments. Although recent tool-use benc

model-releasesarxiv-cs-cl
25 Jun 2026
Agents

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning

DGX agent

arXiv:2605.20342v1 Announce Type: new Abstract: Training large multimodal models (LMMs) via reinforcement learning (RL) to natively invoke video-processing tools (e.g., cropping) has become a promisin

agentsarxiv-cs-cv
21 May 2026
Agents

Track usage and costs across your users, agents and tools. We're also shipping better access controls for your tools & agents: - Set spend l…

DGX agent

Track usage and costs across your users, agents and tools. We're also shipping better access controls for your tools & agents: - Set spend limits on agents - Set spend limits for users - Control what

agentsharrison-chase--x
15 Apr 2026
Safety

Do LLMs Know Tool Irrelevance? Demystifying Structural Alignment Bias in Tool Invocations

DGX agent

arXiv:2604.11322v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated impressive capabilities in utilizing external tools. In practice, however, LLMs are often exposed to to

safetyarxiv-cs-ai
14 Apr 2026
Model Releases

// The Bitter Lesson of Tool Calling // Tool calling is a design choice, and the defaults are quietly costing accuracy. How so? New research…

DGX agent

// The Bitter Lesson of Tool Calling // Tool calling is a design choice, and the defaults are quietly costing accuracy. How so? New research releases a generation-spanning comparison of programmatic t

model-releasesdair-ai--x
10 Aug 2026
Model Releases

Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools

DGX agent

arXiv:2608.04719v1 Announce Type: new Abstract: Agent evaluations tell us that a model picked the wrong tool, but rarely why. We introduce canary tools: diagnostic probe tools planted in an agent's Mo

model-releasesarxiv-cs-ai
6 Aug 2026
Agents

What Does Vision Tool-Use Reinforcement Learning Really Learn? Disentangling Tool-Induced and Intrinsic Effects for Crop-and-Zoom

DGX agent

arXiv:2602.01334v2 Announce Type: replace Abstract: Vision tool-use reinforcement learning (RL) can equip vision language models with visual operators such as crop-and-zoom and achieves strong perform

agentsarxiv-cs-cv
22 May 2026
Model Releases

Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback

DGX agent

arXiv:2605.17453v1 Announce Type: cross Abstract: Tool-using LLM agents increasingly rely on external tools to make consequential decisions, yet most existing agent-security benchmarks and defenses im

model-releasesarxiv-cs-cl
19 May 2026
Agents

Learning to Rewrite Tool Descriptions for Reliable LLM-Agent Tool Use

DGX agent

arXiv:2602.20426v2 Announce Type: replace Abstract: While most efforts to improve LLM-based tool-using agents focus on the agent itself - through larger models, better prompting, or fine-tuning - agen

agentsarxiv-cs-ai
30 Apr 2026
Agents

New in LangSmith Fleet: Tool access controls and usage tracking. 📊 Track cost and usage by user, agent, and tool from a single dashboard 💳…

DGX agent

New in LangSmith Fleet: Tool access controls and usage tracking. 📊 Track cost and usage by user, agent, and tool from a single dashboard 💳 Set spend limits per user or team to prevent surprises ✅ Cont

agentsharrison-chase--x
15 Apr 2026
Agents

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning

DGX agent

arXiv:2608.09682v1 Announce Type: new Abstract: Tool-augmented vision-language models increasingly 'think with images': they call crop, zoom, or code tools and reason over the returned pixels. However

agentsarxiv-cs-cv
11 Aug 2026
Model Releases

A Few Neurons Reveal When LLMs Misuse Tools: Sparse Detection and Selective Steering for Reliable Tool Use

DGX agent

arXiv:2608.00218v1 Announce Type: new Abstract: Agentic LLMs exhibit three consequential tool-use failures: invalid arguments (validity), unnecessary calls (over-calling), and omitted calls when tools

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

ToolChoiceConfusion: Causal Minimal Tool Filtering for Reliable LLM Agents

DGX agent

arXiv:2606.06284v1 Announce Type: new Abstract: Large language model agents increasingly rely on external tools, but larger tool menus can reduce reliability and efficiency by increasing wrong-tool ca

model-releasesarxiv-cs-ai
6 Jun 2026
Agents

JTPRO: A Joint Tool-Prompt Reflective Optimization Framework for Language Agents

DGX agent

arXiv:2604.19821v1 Announce Type: new Abstract: Large language model (LLM) agents augmented with external tools often struggle as number of tools grow large and become domain-specific. In such setting

agentsarxiv-cs-ai
23 Apr 2026
Model Releases

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents

DGX agent

arXiv:2607.04686v1 Announce Type: cross Abstract: Tool calling is central to modern language model agents, but aggregate benchmark scores often hide where tool use fails. A model that never calls a ne

model-releasesarxiv-cs-ai
7 Jul 2026
Agents

Contract2Tool: Learning Preconditions and Effects for Reliable Tool-Augmented LLM Agents

DGX agent

arXiv:2606.07904v1 Announce Type: new Abstract: Tool-augmented large language model agents increasingly rely on external APIs, but standard tool schemas describe how to call a tool, not when the tool

agentsarxiv-cs-ai
9 Jun 2026
Model Releases

How Many Tools Should an LLM Agent See? A Chance-Corrected Answer

DGX agent

arXiv:2605.24660v1 Announce Type: cross Abstract: Before an LLM agent can use a tool, a retrieval system must decide which candidate tools to show to the agent. How long should that shortlist be? Show

model-releasesarxiv-cs-ai
26 May 2026
Safety

TInR: Exploring Tool-Internalized Reasoning in Large Language Models

DGX agent

arXiv:2604.10788v1 Announce Type: cross Abstract: Tool-Integrated Reasoning (TIR) has emerged as a promising direction by extending Large Language Models' (LLMs) capabilities with external tools durin

safetyarxiv-cs-ai
14 Apr 2026
Agents

ToolScope: Enhancing LLM Agent Tool Use through Tool Merging and Context-Aware Filtering

DGX agent

arXiv:2510.20036v2 Announce Type: replace Abstract: Large language model (LLM) agents rely on external tools to solve complex tasks, but real-world toolsets often contain redundant tools with overlapp

agentsarxiv-cs-cl
12 May 2026
Model Releases

Tool Attention Is All You Need: Dynamic Tool Gating and Lazy Schema Loading for Eliminating the MCP/Tools Tax in Scalable Agentic Workflows

DGX agent

arXiv:2604.21816v1 Announce Type: new Abstract: The Model Context Protocol (MCP) has become a common interface for connecting large language model (LLM) agents to external tools, but its reliance on s

model-releasesarxiv-cs-ai
24 Apr 2026
Model Releases

New release of LLM adds support for reasoning traces, OpenAI Responses, server-side tools, and smarter logging

DGX agent

I released LLM 0.32 this morning, the most significant new version of LLM since the initial launch of the project. The new version includes support for visible reasoning traces, server-side provider t

model-releasessimon-willison
4 Aug 2026
Model Releases

Do Multimodal Agents Really Benefit from Tool Use? A Systematic Study of Capability Gains

DGX agent

arXiv:2606.02357v1 Announce Type: cross Abstract: Tool-augmented multimodal agents show strong benchmark gains, often taken as evidence that agents have learned to use tools. We argue that this interp

model-releasesarxiv-cs-ai
2 Jun 2026
Agents

Are Tools All We Need? Unveiling the Tool-Use Tax in LLM Agents

DGX agent

arXiv:2605.00136v1 Announce Type: new Abstract: Tool-augmented reasoning has become a popular direction for LLM-based agents, and it is widely assumed to improve reasoning and reliability. However, we

agentsarxiv-cs-ai
5 May 2026
Agents

☁️ Salesforce tools now in Fleet One of the most requested features we've gotten, and it's now a first-class supported tool in Fleet! Just s…

DGX agent

☁️ Salesforce tools now in Fleet One of the most requested features we've gotten, and it's now a first-class supported tool in Fleet! Just sign in with your Salesforce account, and start using it imme

agentsharrison-chase--x
14 Apr 2026
Agents

FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification

DGX agent

arXiv:2607.28225v1 Announce Type: new Abstract: Agentic vision-language models (VLMs), which interleave textual reasoning with explicit tool calls such as cropping and code-based image manipulation, h

agentsarxiv-cs-cv
31 Jul 2026
Model Releases

Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning

DGX agent

arXiv:2605.19852v1 Announce Type: new Abstract: Tool-augmented reasoning has emerged as a promising direction for enhancing the reasoning capabilities of multimodal large language models (MLLMs). Howe

model-releasesarxiv-cs-cl
20 May 2026
Model Releases

Opened a llama.cpp discussion about whether custom GBNF grammars can compose with tool calls in llama-server. Right now tools work alone, gr…

DGX agent

Opened a llama.cpp discussion about whether custom GBNF grammars can compose with tool calls in llama-server. Right now tools work alone, grammar works alone, but tools+grammar doesn't. If you use lla

model-releasesclem-delangue--x
26 Apr 2026
Model Releases

Control Under Compression: Reliability Frontiers for Tool-Using Agents

DGX agent

arXiv:2608.01056v1 Announce Type: cross Abstract: Tool-using language-model agents are governed not only by task prompts but also by persistent system-side instructions that specify tools, arguments,

model-releasesarxiv-cs-cl
4 Aug 2026
← Previous
1
Next →
9,952 results
← Previous
123…208
Next →