AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
9,953 results
21 Jul 2026

We talked about Claude Code, Claude Tag, Fable, coding agent security, evals, tool design, and how Anthropic use these tools themselves Clau…

Model ReleasesDGX agent

We talked about Claude Code, Claude Tag, Fable, coding agent security, evals, tool design, and how Anthropic use these tools themselves Claude Tag (Claude Code via Slack) is already landing 65% of the

30 Jun 2026

Learning How to Use Tools, Not Just When: Pattern-Aware Tool-Integrated Reasoning

TutorialsDGX agent

arXiv:2509.23292v4 Announce Type: replace Abstract: Tool-integrated reasoning (TIR) has become a key approach for improving large reasoning models (LRMs) on complex problems. Prior work has mainly stu

Entity Binding Failures in Tool-Augmented Agents

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety
DGX agent

arXiv:2606.30531v1 Announce Type: new Abstract: Tool-augmented language-model agents are often evaluated by whether they select the correct tool, produce valid API arguments, and complete the requeste

29 Jun 2026

Finally, a people-search tool that actually works. Most people-search tools sell you a frozen list. @CLODOAI searches the live web, reads th…

ResearchDGX agent

Finally, a people-search tool that actually works. Most people-search tools sell you a frozen list. @CLODOAI searches the live web, reads the signals, and tells you why this specific person, right now

ToolPrivacyBench: Benchmarking Purpose-Bound Privacy in Tool-Using LLM Agents

Model ReleasesDGX agent

arXiv:2606.28061v1 Announce Type: cross Abstract: Large language models (LLMs) have increasingly moved from standalone text generation systems to agents that invoke external tools, access environments

9 Jun 2026

Happy to present that terminal tool calls will be in prettified markdown codeblocks now if you have display tool calls enabled in Hermes Age…

AgentsDGX agent

Nous Research announced an enhancement to their Hermes AI model where terminal tool calls will now be displayed in formatted markdown codeblocks when the display tool calls feature is enabled. This im

Bidirectional Semantic Complementary Tool Retrieval for Remote Sensing Agents

Model ReleasesDGX agent

arXiv:2606.07538v1 Announce Type: cross Abstract: Large language model (LLM)-based agents provide a novel paradigm for the automated processing of remote sensing(RS) data. Their success in complex RS

2 Jun 2026

Diversity Over Frequency: Rethinking Tool Use in Visual Chain-of-Thought Agents

AgentsDGX agent

arXiv:2606.00096v1 Announce Type: cross Abstract: Visual agents employ external visual tools within visual chains of thought to incorporate fine-grained evidence. While prior work has mainly studied t

Learning When Not to Act: Mitigating Tool Abuse in Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2606.02132v1 Announce Type: new Abstract: Agentic reinforcement learning can induce tool abuse, where models overuse external tools even for queries solvable by internal reasoning. Existing appr

12 May 2026

Anthropic announces 12 Claude plugins for the legal sector, including a 'commercial counsel' tool for reviewing vendor agreements and a bar exam study tool (Rachel Metz/Bloomberg)

Model ReleasesDGX agent

Rachel Metz / Bloomberg: Anthropic announces 12 Claude plugins for the legal sector, including a “commercial counsel” tool for reviewing vendor agreements and a bar exam study tool — Anthropic PBC is

LLM Agents Already Know When to Call Tools -- Even Without Reasoning

Model ReleasesDGX agent

arXiv:2605.09252v1 Announce Type: new Abstract: Tool-augmented LLM agents tend to call tools indiscriminately, even when the model can answer directly. Each unnecessary call wastes API fees and latenc

CoCoDA: Co-evolving Compositional DAG for Tool-Augmented Agents

AgentsDGX agent

arXiv:2605.08399v1 Announce Type: new Abstract: Tool-augmented language models can extend small language models with external executable skills, but scaling the tool library creates a coupled challeng

11 May 2026

Tool Calling is Linearly Readable and Steerable in Language Models

Model ReleasesDGX agent

arXiv:2605.07990v1 Announce Type: cross Abstract: When a tool-calling agent picks the wrong tool, the failure is invisible until execution: the email gets sent, the meeting gets missed. Probing 12 ins

OrchJail: Jailbreaking Tool-Calling Text-to-Image Agents by Orchestration-Guided Fuzzing

SafetyDGX agent

arXiv:2605.07414v1 Announce Type: cross Abstract: Tool-calling text-to-image (T2I) agents can plan and execute multi-step tool chains to accomplish complex generation and editing queries. However, thi

7 May 2026

I'm really excited about this as a new tool in our interpretability tool kit

SafetyDGX agent

I'm really excited about this as a new tool in our interpretability tool kit In a new paper, we present NLAs, an unsupervised method for converting an LLM's internal state into human-readable text. I'

20 Apr 2026

Dynamic Tool Dependency Retrieval for Lightweight Function Calling

Model ReleasesDGX agent

arXiv:2512.17052v4 Announce Type: replace Abstract: Function calling agents powered by Large Language Models (LLMs) select external tools to automate complex tasks. On-device agents typically use a re

GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows

Model ReleasesDGX agent

arXiv:2604.15715v1 Announce Type: cross Abstract: The development of general-purpose agents requires a shift from executing simple instructions to completing complex, real-world productivity workflows

18 Apr 2026

I prefer my design tool to be more closely integrated with where my agents work. I spent a few hours building my own design tool (inspired b…

Model ReleasesDGX agent

I prefer my design tool to be more closely integrated with where my agents work. I spent a few hours building my own design tool (inspired by Claude Design) inside my orchestrator. I can use this with

12 Aug 2026

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs

Model ReleasesDGX agent

arXiv:2608.10042v1 Announce Type: cross Abstract: Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool u

Continuous Interaction Diffusion: A Diffusion-Native Runtime for Asynchronous Tool-Augmented Reasoning

AgentsDGX agent

arXiv:2608.10438v1 Announce Type: new Abstract: Large language models increasingly rely on external tools to access up-to-date information, perform computation, and interact with the outside world. Fo

11 Aug 2026

OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories

Model ReleasesDGX agent

arXiv:2608.08557v1 Announce Type: new Abstract: Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding. The prevailing

ToolVision: Learning When and How to Use Visual Tools with Capability-Aligned Supervision

AgentsDGX agent

arXiv:2608.08907v1 Announce Type: cross Abstract: Thinking with images allows a multimodal model to compensate for limited perception by invoking visual tools through code. Yet the prevailing SFT-then

8 Jul 2026

Controlling Tool Use with Heading-Specific Activation Steering

SafetyDGX agent

arXiv:2607.05790v1 Announce Type: new Abstract: Tool-augmented large language models extend their capabilities beyond parametric knowledge through external tools, but tend to invoke them unnecessarily

When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents

SafetyDGX agent

arXiv:2606.20023v2 Announce Type: replace-cross Abstract: As LLM agents increasingly select tools autonomously, their choices among tools with different privileges become safety-relevant. However, pri

4 Jul 2026

Better Models: Worse Tools

Model ReleasesDGX agent

Better Models: Worse Tools Armin reports on a weird problem he ran into while hacking on Pi: The short version is that newer Claude models sometimes call Pi’s edit tool with extra, invented fields in

8 Jun 2026

NTILC: Neural Tool Invocation via Learned Compression

Model ReleasesDGX agent

arXiv:2606.06566v1 Announce Type: cross Abstract: Agentic tool-calling language models depend on large registries of callable APIs, functions, and local actions. Placing full tool specifications direc

29 May 2026

ParaTool: Shifting Tool Representations from Context to Parameters

Model ReleasesDGX agent

arXiv:2605.29561v1 Announce Type: new Abstract: Tool calling extends large language models (LLMs) by enabling grounded interaction with external executable interfaces, thereby supporting environment-c

19 May 2026

Internalizing Tool Knowledge in Small Language Models via QLoRA Fine-Tuning

Model ReleasesDGX agent

arXiv:2605.17774v1 Announce Type: new Abstract: Large language models are increasingly used as planning components in agentic systems, but current tool-use pipelines often require full tool schemas to

10 Apr 2026

Tool-MCoT: Tool Augmented Multimodal Chain-of-Thought for Content Safety Moderation

SafetyDGX agent

arXiv:2604.06205v1 Announce Type: cross Abstract: The growth of online platforms and user content requires strong content moderation systems that can handle complex inputs from various media types. Wh

@hwchase17 Ngl I really like this direction. The more AGENTS.md, skills, and tool config start looking like portable interfaces instead of a…

AgentsDGX agent

A developer expressed enthusiasm for the emerging convergence of `AGENTS.md`, agent skills (SKILL.md), and tool configuration toward portable, cross-tool interfaces rather than siloed, tool-specifi...

5 Aug 2026

Big new release of my LLM CLI tool and Python library for talking to hundreds of different LLMs - reasoning traces, OpenAI Responses support…

ToolsDGX agent

Big new release of my LLM CLI tool and Python library for talking to hundreds of different LLMs - reasoning traces, OpenAI Responses support, server-side tools, smarter logging and a whole lot more ht

Getting the Parameters Right: A Difficulty-Graded Benchmark and Probe-Guided Training for LLM Tool Calls

Model ReleasesDGX agent

arXiv:2608.03071v1 Announce Type: new Abstract: Large language model agents derive much of their capability from tool use. Existing research on tool use has largely focused on selecting the right tool

4 Aug 2026

VC-Tooler: Learning Compositional and Adaptive Visual Tool Use

AgentsDGX agent

arXiv:2608.02217v1 Announce Type: new Abstract: Agentic multimodal reasoning extends passive image understanding by allowing VLMs to actively acquire and refine visual evidence through visual tool int

29 Jul 2026

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing

Local AiDGX agent

arXiv:2607.25993v1 Announce Type: new Abstract: Ultra-high-resolution (UHR) remote-sensing (RS) imagery provides fine-grained Earth-observation evidence over city-scale scenes, but poses a fundamental

9 Jul 2026

MCP tool design: Practical approaches and tradeoffs

AgentsDGX agent

MCP tool design involves deciding the granularity of tools—whether they map to individual API calls or complete workflows—which directly impacts how many tools agents need and their effectiveness. Eff

7 Jul 2026

Securing Multi-Tool AI Agent Chains With Dynamic, Real-Time Compositional Policies

SafetyDGX agent

arXiv:2607.03423v1 Announce Type: cross Abstract: Modern AI agent implementations such as frontier coding agents chain multiple tools at runtime that create a security surface that per-tool guardrails

13 May 2026

GRAFT: Graph-Tokenized LLMs for Tool Planning

SafetyDGX agent

arXiv:2605.11706v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to complete complex tasks by selecting and coordinating external tools across multiple steps. This re

29 Apr 2026

LLM 0.32a0 is a major backwards-compatible refactor

Model ReleasesDGX agent

I just released LLM 0.32a0, an alpha release of my LLM Python library and CLI tool for accessing LLMs, with some consequential changes that I've been working towards for quite a while. Previous versio

23 Apr 2026

Meta-Tool: Efficient Few-Shot Tool Adaptation for Small Language Models

Model ReleasesDGX agent

arXiv:2604.20148v1 Announce Type: cross Abstract: Can small language models achieve strong tool-use performance without complex adaptation mechanisms? This paper investigates this question through Met

Visual Reasoning through Tool-supervised Reinforcement Learning

AgentsDGX agent

arXiv:2604.19945v1 Announce Type: new Abstract: In this paper, we investigate the problem of how to effectively master tool-use to solve complex visual reasoning tasks for Multimodal Large Language Mo

SceneOrchestra: Efficient Agentic 3D Scene Synthesis via Full Tool-Call Trajectory Generation

Model ReleasesDGX agent

arXiv:2604.19907v1 Announce Type: new Abstract: Recent agentic frameworks for 3D scene synthesis have advanced realism and diversity by integrating heterogeneous generation and editing tools. These to

8 Apr 2026

PS: I finally got around to trying out @randal_olson 's Tufte Test tool to prettify the benchmark plot. Great tool 👌! https://www.goodeyela…

Model ReleasesDGX agent

Sebastian Raschka (rasbt) used Randal Olson's Tufte Test tool, developed by Goodeye Labs, to improve the visual quality of a machine learning benchmark plot. The Tufte Test encodes seven of Tufte'...

3 Aug 2026

CAGE: Certified Authorization under Typed-Return Uncertainty for Tool-Using Agents

SafetyDGX agent

arXiv:2607.29190v1 Announce Type: new Abstract: Tool-using LLM agents act on typed tool returns, records pairing provenance and categorical fields with numerical values. Runtime permission gates gener

28 May 2026

SynthTools: A Framework for Scaling Synthetic Tools for Agent Development

AgentsDGX agent

arXiv:2511.09572v2 Announce Type: replace Abstract: For agentic systems to use external tools to solve complex, long-horizon tasks, we need a large set of diverse and controllable tool-use environment

27 May 2026

Enabling Extensible Embodied Capabilities with Tools

SafetyDGX agent

arXiv:2605.26637v1 Announce Type: new Abstract: Most existing embodied intelligence methods formulate perception, reasoning, planning, and control within a unified parameterized policy. Yet these capa

Tool-Schema Compression Enables Agentic RAG Under Constrained Context Budgets

Model ReleasesDGX agent

arXiv:2605.26165v1 Announce Type: cross Abstract: Agentic RAG systems that equip language models with dozens to hundreds of tool definitions face a critical resource conflict: tool schemas consume the

22 May 2026

NaviAgent: Graph-Driven Bilevel Planning for Scalable Tool Orchestration

SafetyDGX agent

arXiv:2506.19500v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) increasingly act as function-call agents that invoke external tools to tackle tasks beyond their static knowledge

5 May 2026

To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling

AgentsDGX agent

arXiv:2605.00737v1 Announce Type: new Abstract: Agentic AI architectures augment LLMs with external tools, unlocking strong capabilities. However, tool use is not always beneficial; some calls may be

toodles from mickey mouse clubhouse was weirdly ahead of its time wake phrase: mickey mouse clubhouse launched in 2006, and “oh toodles” tra…

Model ReleasesDGX agent

toodles from mickey mouse clubhouse was weirdly ahead of its time wake phrase: mickey mouse clubhouse launched in 2006, and “oh toodles” trained toddlers on the assistant wake phrase years before siri

4 May 2026

PORTool: Importance-Aware Policy Optimization with Rewarded Tree for Multi-Tool-Integrated Reasoning

SafetyDGX agent

arXiv:2510.26020v2 Announce Type: replace Abstract: Multi-tool-integrated reasoning enables LLM-empowered tool-use agents to solve complex tasks by interleaving natural-language reasoning with calls t

21 Apr 2026

I for one would be delighted to see OpenAI commit to maintaining a tool like this in the long-term, I'm already nervous about mine going sta…

ToolsDGX agent

Simon Willison expresses hope that OpenAI will commit to long-term maintenance of an AI tool, citing concerns about his own tool potentially becoming stale or discontinued. The post reflects broader u

17 Apr 2026

One portal, unlimited possibilities. You can now access Modal via Tool Gateway by @NousResearch, makers of Hermes Agent. Check it out 👇

AgentsDGX agent

One portal, unlimited possibilities. You can now access Modal via Tool Gateway by @NousResearch, makers of Hermes Agent. Check it out 👇 Tool Gateway is now live in Nous Portal. No separate accounts, n

15 Apr 2026

Can AI Tools Transform Low-Demand Math Tasks? An Evaluation of Task Modification Capabilities

Model ReleasesDGX agent

arXiv:2604.12743v1 Announce Type: new Abstract: While recent research has explored AI tools' ability to classify the quality of mathematical tasks (arXiv:2603.03512), little is known about their capac

11 Apr 2026

All Tools

ConceptsDGX agent

Auto-generated index of all tools mentioned across the wiki.

28 Jul 2026

Intent-Governed Tool Authorization for AI Agents

SafetyDGX agent

arXiv:2606.22916v2 Announce Type: replace Abstract: AI agents increasingly act through external tools: they read private data, construct structured payloads, submit write requests, export records, and

25 Jun 2026

Constraint Tax in Open-Weight LLMs: An Empirical Study of Tool Calling Suppression Under Structured Output Constraints

Model ReleasesDGX agent

arXiv:2606.25605v1 Announce Type: new Abstract: Tool Calling and Structured Output are two core capabilities of modern Agent systems, yet their interaction under joint deployment conditions remains in

4 Jun 2026

EVA-Bench Data 2.0: 3 Domains, 121 Tools, 213 Scenarios

ToolsDGX agent

EVA-Bench Data 2.0 is an expanded benchmark dataset containing tools and scenarios across 3 domains, featuring 121 tools and 213 test scenarios for evaluating AI agent performance. This dataset enable

3 Jun 2026

Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2606.03762v1 Announce Type: cross Abstract: Agentic reinforcement learning (RL) equips large language models (LLMs) with tool-use capabilities that substantially improve reasoning on complex tas

14 May 2026

RS-Claw: Progressive Active Tool Exploration via Hierarchical Skill Trees for Remote Sensing Agents

Model ReleasesDGX agent

arXiv:2605.13391v1 Announce Type: new Abstract: The rise of multi-modal large language models (MLLMs) is shifting remote sensing (RS) intelligence from 'see' to 'action', as OpenClaw-style frameworks

14 Apr 2026

Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning

AgentsDGX agent

arXiv:2604.09813v1 Announce Type: new Abstract: Existing synthetic tool-use corpora are primarily designed for offline supervised fine-tuning, yet reinforcement learning (RL) requires executable envir

← Previous
1234…166
Next →