AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlog
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
9,948 results
Local Ai

A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding

DGX agent

arXiv:2512.21414v2 Announce Type: replace Abstract: Recent tool-use frameworks powered by vision-language models (VLMs) improve image understanding by grounding model predictions with specialized tool

local-aiarxiv-cs-cv
10 Jul 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Agents

Ghost Tool Calls: Issue-Time Privacy for Speculative Agent Tools

DGX agent

arXiv:2606.02483v1 Announce Type: cross Abstract: Tool-augmented language agents speculatively issue likely future tool calls to hide latency, but those calls leak inferred user intent to external ser

agentsarxiv-cs-ai
2 Jun 2026
Agents

ARMOR: An Agentic Framework for Reaction Feasibility Prediction via Adaptive Utility-aware Multi-tool Reasoning

DGX agent

arXiv:2605.07103v1 Announce Type: new Abstract: Reaction feasibility prediction, as a fundamental problem in computational chemistry, has benefited from diverse tools enabled by recent advances in art

agentsarxiv-cs-ai
11 May 2026
Model Releases

From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models

DGX agent

arXiv:2511.10899v2 Announce Type: replace Abstract: Tool-augmented Language Models (TaLMs) can invoke external tools to solve problems beyond their parametric capacity. However, it remains unclear whe

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

ToolOmni: Enabling Open-World Tool Use via Agentic learning with Proactive Retrieval and Grounded Execution

DGX agent

arXiv:2604.13787v1 Announce Type: new Abstract: Large Language Models (LLMs) enhance their problem-solving capability by utilizing external tools. However, in open-world scenarios with massive and evo

model-releasesarxiv-cs-cl
16 Apr 2026
Agents

Amplitude-Only FFN Intervention for Tool-Structured LLM Inference Method: Gated Evaluation Protocol, and Cross-Model Empirical Results

DGX agent

arXiv:2607.11183v2 Announce Type: replace Abstract: Large language models increasingly operate as tool-using agents, where small format, argument, or function-call errors can invalidate otherwise plau

agentsarxiv-cs-cl
15 Jul 2026
Industry

The US FDA drops an enforcement complaint against Whoop over its blood pressure tracking tool, reversing a July 2025 warning letter; Whoop is updating the tool (Samantha Kelly/Bloomberg)

DGX agent

Samantha Kelly / Bloomberg: The US FDA drops an enforcement complaint against Whoop over its blood pressure tracking tool, reversing a July 2025 warning letter; Whoop is updating the tool — The US Foo

industrytechmeme
24 Jun 2026
Safety

Implicit Hierarchical GRPO: Decoupling Tool Invocation from Execution for Tool-Integrated Mathematical Reasoning

DGX agent

arXiv:2605.18500v1 Announce Type: new Abstract: Large language models (LLMs) have increasingly leveraged tool invocation to enhance their reasoning capabilities. However, existing approaches typically

safetyarxiv-cs-cl
19 May 2026
Model Releases

TRACER: Verifiable Generative Provenance for Multimodal Tool-Using Agents

DGX agent

arXiv:2605.09934v1 Announce Type: new Abstract: Multimodal large language models increasingly solve vision-centric tasks by calling external tools for visual inspection, OCR, retrieval, calculation, a

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Beyond the Query: 5 Scenarios Laying the Foundation for the Agentic Era

DGX agent

Accessing enterprise data is shifting from static reports to dynamic use by autonomous systems. To keep up, organizations must route fragmented data from SaaS, IoT, and legacy sources into secure, sca

model-releasesgoogle-cloud-ai
18 May 2026
Model Releases

The Bitter Lesson of Tool Calling

DGX agent

arXiv:2608.06370v1 Announce Type: new Abstract: Tool use transforms LLMs into agents that act beyond their training data, and for code-capable models, programmatic tool calling extends this further by

model-releasesarxiv-cs-cl
7 Aug 2026
Model Releases

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise

DGX agent

arXiv:2608.02372v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in task-oriented dialogue systems that support multi-step decision-making in high-stakes domains

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

We talked about Claude Code, Claude Tag, Fable, coding agent security, evals, tool design, and how Anthropic use these tools themselves Clau…

DGX agent

We talked about Claude Code, Claude Tag, Fable, coding agent security, evals, tool design, and how Anthropic use these tools themselves Claude Tag (Claude Code via Slack) is already landing 65% of the

model-releasessimon-willison--x
21 Jul 2026
Tutorials

Learning How to Use Tools, Not Just When: Pattern-Aware Tool-Integrated Reasoning

DGX agent

arXiv:2509.23292v4 Announce Type: replace Abstract: Tool-integrated reasoning (TIR) has become a key approach for improving large reasoning models (LRMs) on complex problems. Prior work has mainly stu

tutorialsarxiv-cs-ai
30 Jun 2026
Research

Finally, a people-search tool that actually works. Most people-search tools sell you a frozen list. @CLODOAI searches the live web, reads th…

DGX agent

Finally, a people-search tool that actually works. Most people-search tools sell you a frozen list. @CLODOAI searches the live web, reads the signals, and tells you why this specific person, right now

researchdair-ai--x
29 Jun 2026
Agents

Happy to present that terminal tool calls will be in prettified markdown codeblocks now if you have display tool calls enabled in Hermes Age…

DGX agent

Nous Research announced an enhancement to their Hermes AI model where terminal tool calls will now be displayed in formatted markdown codeblocks when the display tool calls feature is enabled. This im

agentsnous-research--x
9 Jun 2026
Agents

Diversity Over Frequency: Rethinking Tool Use in Visual Chain-of-Thought Agents

DGX agent

arXiv:2606.00096v1 Announce Type: cross Abstract: Visual agents employ external visual tools within visual chains of thought to incorporate fine-grained evidence. While prior work has mainly studied t

agentsarxiv-cs-ai
2 Jun 2026
Model Releases

Anthropic announces 12 Claude plugins for the legal sector, including a 'commercial counsel' tool for reviewing vendor agreements and a bar exam study tool (Rachel Metz/Bloomberg)

DGX agent

Rachel Metz / Bloomberg: Anthropic announces 12 Claude plugins for the legal sector, including a “commercial counsel” tool for reviewing vendor agreements and a bar exam study tool — Anthropic PBC is

model-releasestechmeme
12 May 2026
Model Releases

LLM Agents Already Know When to Call Tools -- Even Without Reasoning

DGX agent

arXiv:2605.09252v1 Announce Type: new Abstract: Tool-augmented LLM agents tend to call tools indiscriminately, even when the model can answer directly. Each unnecessary call wastes API fees and latenc

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Tool Calling is Linearly Readable and Steerable in Language Models

DGX agent

arXiv:2605.07990v1 Announce Type: cross Abstract: When a tool-calling agent picks the wrong tool, the failure is invisible until execution: the email gets sent, the meeting gets missed. Probing 12 ins

model-releasesarxiv-cs-ai
11 May 2026
Safety

I'm really excited about this as a new tool in our interpretability tool kit

DGX agent

I'm really excited about this as a new tool in our interpretability tool kit In a new paper, we present NLAs, an unsupervised method for converting an LLM's internal state into human-readable text. I'

safetyjan-leike--x
7 May 2026
Model Releases

Dynamic Tool Dependency Retrieval for Lightweight Function Calling

DGX agent

arXiv:2512.17052v4 Announce Type: replace Abstract: Function calling agents powered by Large Language Models (LLMs) select external tools to automate complex tasks. On-device agents typically use a re

model-releasesarxiv-cs-lg
20 Apr 2026
Model Releases

GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows

DGX agent

arXiv:2604.15715v1 Announce Type: cross Abstract: The development of general-purpose agents requires a shift from executing simple instructions to completing complex, real-world productivity workflows

model-releasesarxiv-cs-ai
20 Apr 2026
Model Releases

I prefer my design tool to be more closely integrated with where my agents work. I spent a few hours building my own design tool (inspired b…

DGX agent

I prefer my design tool to be more closely integrated with where my agents work. I spent a few hours building my own design tool (inspired by Claude Design) inside my orchestrator. I can use this with

model-releasesdair-ai--x
18 Apr 2026
Model Releases

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs

DGX agent

arXiv:2608.10042v1 Announce Type: cross Abstract: Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool u

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories

DGX agent

arXiv:2608.08557v1 Announce Type: new Abstract: Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding. The prevailing

model-releasesarxiv-cs-cl
11 Aug 2026
Safety

Controlling Tool Use with Heading-Specific Activation Steering

DGX agent

arXiv:2607.05790v1 Announce Type: new Abstract: Tool-augmented large language models extend their capabilities beyond parametric knowledge through external tools, but tend to invoke them unnecessarily

safetyarxiv-cs-ai
8 Jul 2026
Safety

When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents

DGX agent

arXiv:2606.20023v2 Announce Type: replace-cross Abstract: As LLM agents increasingly select tools autonomously, their choices among tools with different privileges become safety-relevant. However, pri

safetyarxiv-cs-ai
8 Jul 2026
Model Releases

Better Models: Worse Tools

DGX agent

Better Models: Worse Tools Armin reports on a weird problem he ran into while hacking on Pi: The short version is that newer Claude models sometimes call Pi’s edit tool with extra, invented fields in

model-releasessimon-willison
4 Jul 2026
Model Releases

NTILC: Neural Tool Invocation via Learned Compression

DGX agent

arXiv:2606.06566v1 Announce Type: cross Abstract: Agentic tool-calling language models depend on large registries of callable APIs, functions, and local actions. Placing full tool specifications direc

model-releasesarxiv-cs-ai
8 Jun 2026
Safety

Learning When Not to Act: Mitigating Tool Abuse in Agentic Reinforcement Learning

DGX agent

arXiv:2606.02132v1 Announce Type: new Abstract: Agentic reinforcement learning can induce tool abuse, where models overuse external tools even for queries solvable by internal reasoning. Existing appr

safetyarxiv-cs-ai
2 Jun 2026
Model Releases

ParaTool: Shifting Tool Representations from Context to Parameters

DGX agent

arXiv:2605.29561v1 Announce Type: new Abstract: Tool calling extends large language models (LLMs) by enabling grounded interaction with external executable interfaces, thereby supporting environment-c

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Internalizing Tool Knowledge in Small Language Models via QLoRA Fine-Tuning

DGX agent

arXiv:2605.17774v1 Announce Type: new Abstract: Large language models are increasingly used as planning components in agentic systems, but current tool-use pipelines often require full tool schemas to

model-releasesarxiv-cs-cl
19 May 2026
Safety

Tool-MCoT: Tool Augmented Multimodal Chain-of-Thought for Content Safety Moderation

DGX agent

arXiv:2604.06205v1 Announce Type: cross Abstract: The growth of online platforms and user content requires strong content moderation systems that can handle complex inputs from various media types. Wh

safetyarxiv-cs-ai
10 Apr 2026
Tools

Big new release of my LLM CLI tool and Python library for talking to hundreds of different LLMs - reasoning traces, OpenAI Responses support…

DGX agent

Big new release of my LLM CLI tool and Python library for talking to hundreds of different LLMs - reasoning traces, OpenAI Responses support, server-side tools, smarter logging and a whole lot more ht

toolssimon-willison--x
5 Aug 2026
Agents

VC-Tooler: Learning Compositional and Adaptive Visual Tool Use

DGX agent

arXiv:2608.02217v1 Announce Type: new Abstract: Agentic multimodal reasoning extends passive image understanding by allowing VLMs to actively acquire and refine visual evidence through visual tool int

agentsarxiv-cs-cv
4 Aug 2026
Local Ai

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing

DGX agent

arXiv:2607.25993v1 Announce Type: new Abstract: Ultra-high-resolution (UHR) remote-sensing (RS) imagery provides fine-grained Earth-observation evidence over city-scale scenes, but poses a fundamental

local-aiarxiv-cs-cv
29 Jul 2026
Agents

MCP tool design: Practical approaches and tradeoffs

DGX agent

MCP tool design involves deciding the granularity of tools—whether they map to individual API calls or complete workflows—which directly impacts how many tools agents need and their effectiveness. Eff

agentsaws-ml-blog
9 Jul 2026
Safety

Securing Multi-Tool AI Agent Chains With Dynamic, Real-Time Compositional Policies

DGX agent

arXiv:2607.03423v1 Announce Type: cross Abstract: Modern AI agent implementations such as frontier coding agents chain multiple tools at runtime that create a security surface that per-tool guardrails

safetyarxiv-cs-ai
7 Jul 2026
Safety

GRAFT: Graph-Tokenized LLMs for Tool Planning

DGX agent

arXiv:2605.11706v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to complete complex tasks by selecting and coordinating external tools across multiple steps. This re

safetyarxiv-cs-lg
13 May 2026
Model Releases

LLM 0.32a0 is a major backwards-compatible refactor

DGX agent

I just released LLM 0.32a0, an alpha release of my LLM Python library and CLI tool for accessing LLMs, with some consequential changes that I've been working towards for quite a while. Previous versio

model-releasessimon-willison
29 Apr 2026
Model Releases

Meta-Tool: Efficient Few-Shot Tool Adaptation for Small Language Models

DGX agent

arXiv:2604.20148v1 Announce Type: cross Abstract: Can small language models achieve strong tool-use performance without complex adaptation mechanisms? This paper investigates this question through Met

model-releasesarxiv-cs-ai
23 Apr 2026
Agents

Visual Reasoning through Tool-supervised Reinforcement Learning

DGX agent

arXiv:2604.19945v1 Announce Type: new Abstract: In this paper, we investigate the problem of how to effectively master tool-use to solve complex visual reasoning tasks for Multimodal Large Language Mo

agentsarxiv-cs-cv
23 Apr 2026
Agents

@hwchase17 Ngl I really like this direction. The more AGENTS.md, skills, and tool config start looking like portable interfaces instead of a…

DGX agent

A developer expressed enthusiasm for the emerging convergence of `AGENTS.md`, agent skills (SKILL.md), and tool configuration toward portable, cross-tool interfaces rather than siloed, tool-specifi...

agentsharrison-chase--x
10 Apr 2026
Model Releases

PS: I finally got around to trying out @randal_olson 's Tufte Test tool to prettify the benchmark plot. Great tool 👌! https://www.goodeyela…

DGX agent

Sebastian Raschka (rasbt) used Randal Olson's Tufte Test tool, developed by Goodeye Labs, to improve the visual quality of a machine learning benchmark plot. The Tufte Test encodes seven of Tufte'...

model-releasessebastian-raschka--x
8 Apr 2026
Agents

ToolVision: Learning When and How to Use Visual Tools with Capability-Aligned Supervision

DGX agent

arXiv:2608.08907v1 Announce Type: cross Abstract: Thinking with images allows a multimodal model to compensate for limited perception by invoking visual tools through code. Yet the prevailing SFT-then

agentsarxiv-cs-ai
11 Aug 2026
Model Releases

Getting the Parameters Right: A Difficulty-Graded Benchmark and Probe-Guided Training for LLM Tool Calls

DGX agent

arXiv:2608.03071v1 Announce Type: new Abstract: Large language model agents derive much of their capability from tool use. Existing research on tool use has largely focused on selecting the right tool

model-releasesarxiv-cs-ai
5 Aug 2026
Safety

CAGE: Certified Authorization under Typed-Return Uncertainty for Tool-Using Agents

DGX agent

arXiv:2607.29190v1 Announce Type: new Abstract: Tool-using LLM agents act on typed tool returns, records pairing provenance and categorical fields with numerical values. Runtime permission gates gener

safetyarxiv-cs-ai
3 Aug 2026
← Previous
1234…208
Next →