AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlog
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
9,953 results
Model Releases

Tool Retrieval Bridge: Aligning Vague Instructions with Retriever Preferences via Bridge Model

DGX agent

arXiv:2604.07816v1 Announce Type: new Abstract: Tool learning has emerged as a promising paradigm for large language models (LLMs) to address real-world challenges. Due to the extensive and irregularl

model-releasesarxiv-cs-cl
10 Apr 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks

DGX agent

arXiv:2608.10357v1 Announce Type: cross Abstract: Long-horizon tool-using agents must reason over user goals, domain policies, tool calls, simulator state, and delayed verifiable rewards. Reinforcemen

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

When History Lies: Evaluating and Improving Tool Use under Misleading Multi-Turn Histories

DGX agent

arXiv:2608.06057v1 Announce Type: new Abstract: Tool-calling agents infer task state from accumulated dialogue and tool traces. In persistent interactions, however, historical traces may remain struct

model-releasesarxiv-cs-ai
7 Aug 2026
Agents

Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance

DGX agent

arXiv:2608.03403v1 Announce Type: new Abstract: The performance bottleneck of agents is increasingly shifting from model capability to the robustness of their execution processes. Tools play a central

agentsarxiv-cs-ai
5 Aug 2026
Agents

Verified Tool Calls Improve LLM Agent Reliability Under Non-Atomic Failures

DGX agent

arXiv:2608.02645v1 Announce Type: cross Abstract: Large Language Model (LLM) agents rely on external tools to perform multistage tasks. Existing agent frameworks typically assume that tool calls are a

agentsarxiv-cs-ai
5 Aug 2026
Model Releases

20 questions for the Agentic Enterprise (and how Agent Platform can help)

DGX agent

If you’re an IT leader, you might be getting a lot of questions about how to build and deploy agents. The pressure to move fast is intense, but the engineering reality is incredibly complex. Where do

model-releasesgoogle-cloud-ai
7 Jul 2026
Model Releases

Determinants and Limits of LLM Security-Tool Orchestration: A Study with HexStrike-AI

DGX agent

arXiv:2607.02873v1 Announce Type: cross Abstract: Large language model agents driving security tool suites over the Model Context Protocol are increasingly common. Yet the factors that bound their cap

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

Governed MCP: Kernel-Level Tool Governance for AI Agents via Logit-Based Safety Primitives

DGX agent

arXiv:2604.16870v2 Announce Type: replace-cross Abstract: AI agents increasingly call external tools (file system, network, APIs) through the Model Context Protocol (MCP). These tool calls are the age

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

Expert Evaluation of Clinical AI Tools on Real Point-of-Care Clinical Queries

DGX agent

arXiv:2606.28960v1 Announce Type: new Abstract: Physicians now pose millions of clinical questions to AI tools each week, yet these tools are evaluated largely on hypothetical or exam-style questions,

model-releasesarxiv-cs-ai
30 Jun 2026
Agents

GROW^2: Grounding Which and Where for Robot Tool Use

DGX agent

arXiv:2606.30632v1 Announce Type: cross Abstract: Can the robot use a plate to cut a cake if no knife is available? Tool use greatly expands robot capabilities, but to use tools creatively beyond thei

agentsarxiv-cs-ai
30 Jun 2026
Research

Geometric Reconstruction of Extrinsic Contact Trajectories using Tactile Sensing and Proprioception for Tool Manipulation

DGX agent

arXiv:2606.22251v1 Announce Type: new Abstract: Tactile sensing enables robots to perceive rich contact information at the grasp, supporting tasks such as object recognition, in-hand pose estimation,

researcharxiv-cs-ro
23 Jun 2026
Model Releases

MedCTA: A Benchmark for Clinical Tool Agents

DGX agent

arXiv:2606.11702v1 Announce Type: cross Abstract: To make clinically grounded decisions, medical AI agents are expected to go beyond simple recognition and be capable of tool retrieval, evidence acqui

model-releasesarxiv-cs-ai
11 Jun 2026
Model Releases

Beyond APIs: Probing the Limits of MLLMs in Physical Tool Use

DGX agent

arXiv:2606.10803v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) excel at utilizing digital APIs and increasingly serve as the 'brain' of embodied AI, instructing robots to i

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

The Impact of Configuring Agentic AI Coding Tools on Build-vs-Buy Decisions: A Study Protocol

DGX agent

arXiv:2606.03907v1 Announce Type: cross Abstract: Agentic AI coding tools write code with increasing autonomy and in doing so decide when to import a library and when to implement functionality from s

model-releasesarxiv-cs-ai
3 Jun 2026
Tools

Uber reportedly now caps coding agents at $1,500/month per employee per tool - seems sensible to me, but it's also an interesting hint at th…

DGX agent

Uber reportedly now caps coding agents at $1,500/month per employee per tool - seems sensible to me, but it's also an interesting hint at the value Uber thinks these tools are providing https://simonw

toolssimon-willison--x
3 Jun 2026
Model Releases

Same Payload, Different Channel: Measuring Trust Asymmetry in Tool-Using Language Models

DGX agent

arXiv:2606.00566v1 Announce Type: cross Abstract: As language models take on agentic roles that span calling external APIs, reading tool outputs, and acting on instructions embedded in third-party con

model-releasesarxiv-cs-cl
2 Jun 2026
Agents

Hermes Agent now has Tool Search, so your agent only loads what it needs

DGX agent

Hermes Agent has been updated with a Tool Search feature that enables agents to dynamically identify and load only the tools necessary for a given task, rather than loading all available tools upfront

agentsnous-research--x
29 May 2026
Safety

MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning

DGX agent

arXiv:2605.26154v1 Announce Type: cross Abstract: LLM-driven agents are capable of selecting external tools to complete users' tasks. However, attackers could compromise such process, steering agents

safetyarxiv-cs-ai
27 May 2026
Agents

Attested Tool-Server Admission: A Security Extension to the Model Context Protocol

DGX agent

arXiv:2605.24248v1 Announce Type: cross Abstract: The Model Context Protocol (MCP) standardizes how a large-language-model (LLM) agent and an external tool server exchange messages, but not trust: a h

agentsarxiv-cs-ai
26 May 2026
Model Releases

TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents

DGX agent

arXiv:2605.16909v1 Announce Type: new Abstract: Tool-using agents are increasingly expected to operate across realistic professional workflows, where they must interpret multimodal inputs, coordinate

model-releasesarxiv-cs-ai
19 May 2026
Safety

Quantitative Certification of Agentic Tool Selection

DGX agent

arXiv:2510.03992v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed in agentic systems, where a fundamental task is mapping user intents to relevant extern

safetyarxiv-cs-ai
14 May 2026
Tools

Introducing voice finder — a new tool to quickly find the right voice for your app from over 600+ voices

DGX agent

Together AI launched Voice Finder, a tool designed to help developers quickly select appropriate voices for their applications from a library of over 600 voice options. The tool streamlines the voice

toolstogether-ai-blog
12 May 2026
Model Releases

Beyond the Black Box: Interpretability of Agentic AI Tool Use

DGX agent

arXiv:2605.06890v1 Announce Type: new Abstract: AI agents are promising for high-stakes enterprise workflows, but dependable deployment remains limited because tool-use failures are difficult to diagn

model-releasesarxiv-cs-ai
11 May 2026
Agents

Position: Agent Should Invoke External Tools ONLY When Epistemically Necessary

DGX agent

arXiv:2506.00886v3 Announce Type: replace Abstract: As large language models evolve into tool-augmented agents, a central question remains unresolved: when is external tool use actually justified? Exi

agentsarxiv-cs-ai
7 May 2026
Model Releases

TSCG: Deterministic Tool-Schema Compilation for Agentic LLM Deployments

DGX agent

arXiv:2605.04107v1 Announce Type: cross Abstract: Production agent frameworks (OpenAI Function Calling, Anthropic Tool Use, MCP) transmit tool schemas as JSON, a format designed for machine parsing, n

model-releasesarxiv-cs-cl
7 May 2026
Agents

SCRIBE: Structured Mid-Level Supervision for Tool-Using Language Models

DGX agent

arXiv:2601.03555v2 Announce Type: replace Abstract: Training reliable tool-augmented agents remains a significant challenge, largely due to the difficulty of credit assignment in multi-step reasoning.

agentsarxiv-cs-ai
28 Apr 2026
Tools

My talk at AI Engineer “Every API is a Tool for Agents” is out on YouTube https://youtu.be/YBYUvGOuotE Thanks to @swyx and the @aiDotEnginee…

DGX agent

A talk titled 'Every API is a Tool for Agents' from the AI Engineer conference is now available on YouTube, discussing how APIs can be leveraged as tools for AI agents. The presentation was facilitate

toolsswyx--x
25 Apr 2026
Model Releases

260 things we announced at Google Cloud Next '26 – a recap

DGX agent

Google Cloud Next ‘26 took place this week in Las Vegas, and the energy was incredible as we welcomed over 32,000 leaders, developers, and partners to explore the Agentic Era with us. Across three key

model-releasesgoogle-cloud-ai
24 Apr 2026
Model Releases

Trajectory2Task: Training Robust Tool-Calling Agents with Synthesized Yet Verifiable Data for Complex User Intents

DGX agent

arXiv:2601.20144v3 Announce Type: replace Abstract: Tool-calling agents are increasingly deployed in real-world customer-facing workflows. Yet most studies on tool-calling agents focus on idealized se

model-releasesarxiv-cs-cl
23 Apr 2026
Model Releases

Latent Preference Modeling for Cross-Session Personalized Tool Calling

DGX agent

arXiv:2604.17886v1 Announce Type: new Abstract: Users often omit essential details in their requests to LLM-based agents, resulting in under-specified inputs for tool use. This poses a fundamental cha

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

The Reasoning Trap: How Enhancing LLM Reasoning Amplifies Tool Hallucination

DGX agent

arXiv:2510.22977v2 Announce Type: replace-cross Abstract: Enhancing the reasoning capabilities of Large Language Models (LLMs) is a key strategy for building Agents that 'think then act.' However, rec

model-releasesarxiv-cs-ai
20 Apr 2026
Local Ai

ClawGuard: A Runtime Security Framework for Tool-Augmented LLM Agents Against Indirect Prompt Injection

DGX agent

arXiv:2604.11790v1 Announce Type: cross Abstract: Tool-augmented Large Language Model (LLM) agents have demonstrated impressive capabilities in automating complex, multi-step real-world tasks, yet rem

local-aiarxiv-cs-ai
14 Apr 2026
Model Releases

Benchmarking LLM Tool-Use in the Wild

DGX agent

arXiv:2604.06185v1 Announce Type: cross Abstract: Fulfilling user needs through Large Language Model multi-turn, multi-step tool-use is rarely a straightforward process. Real user interactions are inh

model-releasesarxiv-cs-ai
10 Apr 2026
Safety

Learning to Coordinate Symbolic Tools: LLM Agents for Verified Sum-of-Squares Certificates

DGX agent

arXiv:2608.00326v2 Announce Type: replace Abstract: Tool calling allows large language models (LLMs) to invoke external computation during problem solving, a useful capability in various fields includ

safetyarxiv-cs-ai
11 Aug 2026
Model Releases

Models as Tools: An Agentic Coordination Framework for Unified Multimodal Visual Tracking

DGX agent

arXiv:2608.00847v1 Announce Type: new Abstract: Most current visual trackers adopt a matching-based architecture trained exclusively on tracking datasets, whose performance gains depend heavily on the

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

MagicSelector: Joint Optimization for Agent Tool Selection via Counterfactual Decomposition and Progressive Reranking

DGX agent

arXiv:2607.17751v2 Announce Type: cross Abstract: We present MagicSelector, a joint optimization framework integrating Counterfactual task decomposition, Progressive reranking, and Dynamic Top-K, desi

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

Automate your agent development lifecycle using any coding agent

DGX agent

Welcome to our latest Gemini Enterprise Agent Platform deep dive, a practical walkthrough where we’ll teach you how to build real-world, production-ready agents starting from step 1. If you haven’t al

model-releasesgoogle-cloud-ai
29 Jul 2026
Safety

Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents

DGX agent

arXiv:2607.19449v1 Announce Type: cross Abstract: Evaluation frameworks for tool-augmented LLM agents focus overwhelmingly on capability metrics or explicit tool crashes, leaving silent infrastructure

safetyarxiv-cs-ai
23 Jul 2026
Model Releases

Context-Aware Force Estimation for Deformable Tool Manipulation in Robotic Environmental Swabbing via Few-Shot Continual Adaptation

DGX agent

arXiv:2607.07574v1 Announce Type: new Abstract: Robotic surface swabbing requires sustained interaction between a compliant tool and heterogeneous environments, where accurate estimation of tip-level

model-releasesarxiv-cs-ro
9 Jul 2026
Safety

PORTS: Preference-Optimized Retrievers for Tool Selection with Large Language Models

DGX agent

arXiv:2607.05441v1 Announce Type: cross Abstract: Integrating external tools with Large Language Models (LLMs) has emerged as a promising paradigm for accomplishing complex tasks. Since LLMs still str

safetyarxiv-cs-ai
8 Jul 2026
Safety

ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents

DGX agent

arXiv:2606.31392v1 Announce Type: new Abstract: Tool-augmented vision-language models (VLMs) can solve multimodal, multi-step tasks by calling external tools, yet they remain fragile in practice. Exis

safetyarxiv-cs-ai
1 Jul 2026
Agents

VTOS: Learning to Orchestrate Vision Tools by Co-Searching Solutions and Observers

DGX agent

arXiv:2606.20728v1 Announce Type: new Abstract: Vision foundation tools such as open-vocabulary detectors, segmentation models, and post-processing operators are powerful building blocks for computer

agentsarxiv-cs-cv
23 Jun 2026
Safety

IAPO: Input Attribution-Aware Policy Optimization for Tool Use in Small Multimodal Agents

DGX agent

arXiv:2606.11652v1 Announce Type: new Abstract: This paper investigates reinforcement learning (RL) methods for improving tool-calling capabilities in multimodal small language model (SLM) agents. Whi

safetyarxiv-cs-lg
11 Jun 2026
Agents

Pushing the Limits of LLM Tool Calling via Experiential Knowledge Integration and Activation

DGX agent

arXiv:2606.10875v1 Announce Type: new Abstract: Large language models (LLMs) rely on tool use to act as autonomous agents, yet often fail in multi-step execution due to insufficient tool-related knowl

agentsarxiv-cs-cl
10 Jun 2026
Safety

Capability-Aligned Hierarchical Learning for Tool-Augmented LLMs

DGX agent

arXiv:2606.09371v1 Announce Type: new Abstract: Tool learning enables LLMs to invoke external tools to accomplish tasks. Prior studies have demonstrated the effectiveness of a hierarchical structure:

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

Token costs are why there will be no saas apocalypse / good dev tools are cached intelligence for agents! The popular theory goes: agents ca…

DGX agent

Token costs are why there will be no saas apocalypse / good dev tools are cached intelligence for agents! The popular theory goes: agents can write code, so they'll just rebuild every tool from scratc

model-releasesclem-delangue--x
5 Jun 2026
Agents

Synthesize and Reward -- Reinforcement Learning for Multi-Step Tool Use in Live Environments

DGX agent

arXiv:2606.03892v1 Announce Type: cross Abstract: Training LLMs to orchestrate multi-step tool calls is held back by three coupled obstacles: realistic stateful execution environments are costly to bu

agentsarxiv-cs-ai
3 Jun 2026
Model Releases

Uber Caps Usage of AI Tools Like Claude Code to Manage Costs

DGX agent

Uber Caps Usage of AI Tools Like Claude Code to Manage Costs I wrote the other day about Uber blowing its 2026 AI budget in four months, and how that wasn't particularly surprising given they would ha

model-releasessimon-willison
3 Jun 2026
← Previous
123456…208
Next →