AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
9,953 results
29 Jul 2026

Automate your agent development lifecycle using any coding agent

Model ReleasesDGX agent

Welcome to our latest Gemini Enterprise Agent Platform deep dive, a practical walkthrough where we’ll teach you how to build real-world, production-ready agents starting from step 1. If you haven’t al

23 Jul 2026

Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents

SafetyDGX agent

arXiv:2607.19449v1 Announce Type: cross Abstract: Evaluation frameworks for tool-augmented LLM agents focus overwhelmingly on capability metrics or explicit tool crashes, leaving silent infrastructure

9 Jul 2026

Context-Aware Force Estimation for Deformable Tool Manipulation in Robotic Environmental Swabbing via Few-Shot Continual Adaptation

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2607.07574v1 Announce Type: new Abstract: Robotic surface swabbing requires sustained interaction between a compliant tool and heterogeneous environments, where accurate estimation of tip-level

Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents

Model ReleasesDGX agent

arXiv:2607.07405v1 Announce Type: new Abstract: Tool-using LLM agents can violate the very policies they are deployed to enforce while appearing to complete the task successfully. In policy-permissive

8 Jul 2026

PORTS: Preference-Optimized Retrievers for Tool Selection with Large Language Models

SafetyDGX agent

arXiv:2607.05441v1 Announce Type: cross Abstract: Integrating external tools with Large Language Models (LLMs) has emerged as a promising paradigm for accomplishing complex tasks. Since LLMs still str

When Does Tool Use Increase the Expressive Power of Finite-Precision Recurrent Models?

AgentsDGX agent

arXiv:2607.06155v1 Announce Type: cross Abstract: Modern sequence models are increasingly deployed as agents that interleave token generation with calls to external tools. We give an exact, architectu

1 Jul 2026

ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents

SafetyDGX agent

arXiv:2606.31392v1 Announce Type: new Abstract: Tool-augmented vision-language models (VLMs) can solve multimodal, multi-step tasks by calling external tools, yet they remain fragile in practice. Exis

23 Jun 2026

VTOS: Learning to Orchestrate Vision Tools by Co-Searching Solutions and Observers

AgentsDGX agent

arXiv:2606.20728v1 Announce Type: new Abstract: Vision foundation tools such as open-vocabulary detectors, segmentation models, and post-processing operators are powerful building blocks for computer

11 Jun 2026

IAPO: Input Attribution-Aware Policy Optimization for Tool Use in Small Multimodal Agents

SafetyDGX agent

arXiv:2606.11652v1 Announce Type: new Abstract: This paper investigates reinforcement learning (RL) methods for improving tool-calling capabilities in multimodal small language model (SLM) agents. Whi

10 Jun 2026

Pushing the Limits of LLM Tool Calling via Experiential Knowledge Integration and Activation

AgentsDGX agent

arXiv:2606.10875v1 Announce Type: new Abstract: Large language models (LLMs) rely on tool use to act as autonomous agents, yet often fail in multi-step execution due to insufficient tool-related knowl

Less Context, Better Agents: Efficient Context Engineering for Long-Horizon Tool-Using LLM Agents

Model ReleasesDGX agent

arXiv:2606.10209v1 Announce Type: new Abstract: Large language models deployed as autonomous agents for enterprise workflows face a key challenge: verbose tool responses from enterprise systems can ca

9 Jun 2026

Capability-Aligned Hierarchical Learning for Tool-Augmented LLMs

SafetyDGX agent

arXiv:2606.09371v1 Announce Type: new Abstract: Tool learning enables LLMs to invoke external tools to accomplish tasks. Prior studies have demonstrated the effectiveness of a hierarchical structure:

5 Jun 2026

Token costs are why there will be no saas apocalypse / good dev tools are cached intelligence for agents! The popular theory goes: agents ca…

Model ReleasesDGX agent

Token costs are why there will be no saas apocalypse / good dev tools are cached intelligence for agents! The popular theory goes: agents can write code, so they'll just rebuild every tool from scratc

DexFuture: Hierarchical Future-State Visuomotor Targeting for Bimanual Dexterous Tool Use

SafetyDGX agent

arXiv:2606.05699v1 Announce Type: new Abstract: Bimanual dexterous tool use remains challenging for robots due to high-dimensional hand configurations and complex hand-tool-object dynamics and contact

3 Jun 2026

Synthesize and Reward -- Reinforcement Learning for Multi-Step Tool Use in Live Environments

AgentsDGX agent

arXiv:2606.03892v1 Announce Type: cross Abstract: Training LLMs to orchestrate multi-step tool calls is held back by three coupled obstacles: realistic stateful execution environments are costly to bu

Uber Caps Usage of AI Tools Like Claude Code to Manage Costs

Model ReleasesDGX agent

Uber Caps Usage of AI Tools Like Claude Code to Manage Costs I wrote the other day about Uber blowing its 2026 AI budget in four months, and how that wasn't particularly surprising given they would ha

2 Jun 2026

Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation

Model ReleasesDGX agent

arXiv:2512.16310v3 Announce Type: replace-cross Abstract: LLM-based agents increasingly use multiple external tools to complete complex tasks. We study Tools Orchestration Privacy Risk (TOP-R): an age

VESTA: Visual Exploration with Statistical Tool Agents

Model ReleasesDGX agent

arXiv:2606.00384v1 Announce Type: new Abstract: Fitting quantitative models to data is a central step in scientific workflows, yet it remains one of the least automated. Recent agent-based systems lev

29 May 2026

LogDx-CI: Benchmarking Log Reduction Tools for LLM Root-Cause Diagnosis

Model ReleasesDGX agent

arXiv:2605.28876v1 Announce Type: cross Abstract: CI failure logs are large (median 5k lines, max 200k in this corpus) and noisy. Coding agents that try to debug them depend on an upstream tool to red

27 May 2026

Augment Engineering: A Methodology for Multi-Tool AI Orchestration Across Professional Domains

ApplicationsDGX agent

arXiv:2605.26146v1 Announce Type: cross Abstract: Organizations increasingly deploy separate purpose-built AI tools across professional domains, often hiring domain specialists for each, recreating th

ChainCaps: Composition-Safe Tool-Using Agents via Monotonic Capability Attenuation

SafetyDGX agent

arXiv:2605.26542v1 Announce Type: cross Abstract: Tool-using agents increasingly operate in open-ended deployment environments, where they compose file systems, web APIs, code interpreters, and enterp

26 May 2026

Tool-Call Dependency Structure is Linearly Decodable in LLM Agent Residual Streams

AgentsDGX agent

arXiv:2605.25310v1 Announce Type: new Abstract: Tool-using LLM agents produce trajectories whose calls form a directed dependency graph: earlier tool outputs supply arguments to later calls. Whether t

19 May 2026

Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIs

Model ReleasesDGX agent

arXiv:2605.17558v1 Announce Type: cross Abstract: Training tool-calling agents requires large-scale trajectory data with verifiable labels, yet existing approaches either synthesize environments that

Prompts Don't Protect: Architectural Enforcement via MCP Proxy for LLM Tool Access Control

Model ReleasesDGX agent

arXiv:2605.18414v1 Announce Type: cross Abstract: Large language models increasingly operate as autonomous agents that select and invoke tools from large registries. We identify a critical gap: when u

16 May 2026

Interesting interpretability paper on tool-using agents. The authors probe hidden states and find the model often recognizes it should call …

AgentsDGX agent

Interesting interpretability paper on tool-using agents. The authors probe hidden states and find the model often recognizes it should call a tool, but fails to actually call one. The mismatch ranges

12 May 2026

PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools

SafetyDGX agent

arXiv:2604.01532v2 Announce Type: replace Abstract: LLM agents are beginning to invoke industrial asset-management tools through the Model Context Protocol (MCP), yet whether they can act reliably on

Useful for Exploration, Risky for Precision: Evaluating AI Tools in Academic Research

ResearchDGX agent

arXiv:2605.10125v1 Announce Type: new Abstract: Artificial intelligence (AI) tools are being incorporated into scientific research workflows with the potential to enhance efficiency in tasks such as d

11 May 2026

AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents

Model ReleasesDGX agent

arXiv:2605.07926v1 Announce Type: new Abstract: As LLM-based agents increasingly rely on external tools, it is important to evaluate their ability to sustain tool-grounded reasoning beyond familiar wo

6 May 2026

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers

Model ReleasesDGX agent

arXiv:2602.00933v2 Announce Type: replace-cross Abstract: The Model Context Protocol (MCP) is rapidly becoming the standard interface for Large Language Models (LLMs) to discover and invoke external t

29 Apr 2026

AdaTooler-V: Adaptive Tool-Use for Images and Videos

Model ReleasesDGX agent

arXiv:2512.16918v3 Announce Type: replace Abstract: Recent advances have shown that multimodal large language models (MLLMs) benefit from multimodal interleaved chain-of-thought (CoT) with vision tool

24 Apr 2026

AgenticQwen: Training Small Agentic Language Models with Dual Data Flywheels for Industrial-Scale Tool Use

AgentsDGX agent

arXiv:2604.21590v1 Announce Type: new Abstract: Modern industrial applications increasingly demand language models that act as agents, capable of multi-step reasoning and tool use in real-world settin

Tool Attention Is All You Need

AgentsDGX agent

Tool Attention Is All You Need // Tool Attention Is All You Need // New research proposes a practical fix for the hidden 'MCP tax.' The work introduces a dynamic tool gating mechanism built on an Inte

23 Apr 2026

You can now use OpenAI OAuth to generate images with Hermes Agent! Update Hermes and use the ‘hermes tools’ command to configure the image g…

AgentsDGX agent

You can now use OpenAI OAuth to generate images with Hermes Agent! Update Hermes and use the ‘hermes tools’ command to configure the image generation tool to use OpenAI! Reminder: we have a $26,000 on

21 Apr 2026

OpenAI's new Euphony tool works almost exactly the same way as my Codex transcript viewer https://tools.simonwillison.net/codex-timeline?url…

ToolsDGX agent

OpenAI's new Euphony tool works almost exactly the same way as my Codex transcript viewer https://tools.simonwillison.net/codex-timeline?url=https%3A%2F%2Fgist.githubusercontent.com%2Fsimonw%2Fa9eb599

Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors

SafetyDGX agent

arXiv:2601.15625v2 Announce Type: replace Abstract: Large language models (LLMs) can call tools effectively, yet they remain brittle in multi-turn execution: after a tool-call error, smaller models of

16 Apr 2026

Below are the Docs to help you get started http://hermes-agent.nousresearch.com/docs/user-guide/features/tool-gateway Thank you to our tool …

AgentsDGX agent

Nous Research announced documentation and resources for getting started with Hermes Agent's tool gateway feature, which appears to be a system for managing and integrating external tools within the He

Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions

Model ReleasesDGX agent

arXiv:2509.18847v3 Announce Type: replace-cross Abstract: Tool-augmented large language models (LLMs) are usually trained with supervised imitation or coarse-grained reinforcement learning that optimi

15 Apr 2026

best tool to create animated wallpapers?

IndustryDGX agent

This Reddit thread on r/ChatGPT discusses community recommendations for the best tools to create animated wallpapers, likely covering AI-assisted options such as using ChatGPT alongside image generato

14 Apr 2026

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks

Model ReleasesDGX agent

arXiv:2604.10015v1 Announce Type: new Abstract: Recent studies demonstrate that tool-calling capability enables large language models (LLMs) to interact with external environments for long-horizon fin

Use of AI Tools: Guidelines to Maintain Academic Integrity in Computing Colleges

TutorialsDGX agent

arXiv:2604.11111v1 Announce Type: cross Abstract: The rapid adoption of AI tools such as ChatGPT has significantly transformed academic practices, offering considerable benefits for both students and

12 Apr 2026

@hwchase17 is right, memory creates lock-in. But not just conversations. Tool registries, hooks, agent conventions -> that's harness memory …

AgentsDGX agent

@hwchase17 is right, memory creates lock-in. But not just conversations. Tool registries, hooks, agent conventions -> that's harness memory too. In my setup I define agent tools once in a shared regis

15 Jul 2026

AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP

AgentsDGX agent

arXiv:2607.11098v2 Announce Type: replace-cross Abstract: Tool-using LLM agents are mostly evaluated assuming all tools work. When a tool times out, returns a week-stale value, or has its description

Wiki Lint Report — 2026-07-15

SynthesesDGX agent

Automated lint: 26 errors, 6728 warnings, 3 info

How to Realize Recursively Self-Improving Agents and Personal Singularity: A Goal-, Scope-, Tool-, and Benchmark-Driven Multi-Agent Architecture

Model ReleasesDGX agent

arXiv:2607.12254v1 Announce Type: new Abstract: Large language model (LLM) agents can increasingly plan, use tools, maintain memory, and execute long-horizon tasks. These advances motivate two linked

11 Aug 2026

LAVE: Latent Visual Evidence-Enhanced Planning for Video Tool-use Agents

AgentsDGX agent

arXiv:2608.07585v1 Announce Type: new Abstract: Long-video understanding requires models to efficiently acquire and reuse sparse visual evidence from long and redundant video streams. Recent video too

VTO: Visual Tool Orchestration for Video Anomaly Detection

Model ReleasesDGX agent

arXiv:2608.08219v1 Announce Type: cross Abstract: Video anomaly detection (VAD) is a critical yet challenging task due to the complex and diverse nature of real-world scenarios. Traditional deep learn

10 Aug 2026

Taxonomy-Driven Analysis of Open-Source AI Risk Mitigation Tools

ApplicationsDGX agent

arXiv:2608.07446v1 Announce Type: cross Abstract: Rapid adoption of large language models (LLMs) in enterprise settings has introduced operational, security, and governance risks. As generative AI app

5 Aug 2026

UrbanAgent: A Tool-Augmented Agent for Cross-System Urban Tasks

Model ReleasesDGX agent

arXiv:2608.03018v1 Announce Type: new Abstract: Modern cities rely on an increasing number of digital services to operate, but residents' daily needs are still difficult to meet. Services are fragment

28 Jul 2026

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios

Model ReleasesDGX agent

arXiv:2607.23722v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as agents that interact with stateful environments over multiple steps: gathering hidden informat

10 Jul 2026

Tool-Making and Self-Evolving LLM Agents in Low-Latency Systems

AgentsDGX agent

arXiv:2607.08010v1 Announce Type: new Abstract: Production LLM agents often waste latency and reliability by regenerating code for the same procedural steps on every request. We replace this inference

7 Jul 2026

When Users Are Happy but Agents Are Wrong: Multi-Dimensional Evaluation of Tool-Augmented Dialogue

Model ReleasesDGX agent

arXiv:2510.19186v3 Announce Type: replace Abstract: Evaluating conversational AI systems that use external tools is challenging, as errors can arise from complex interactions among user, agent, and to

2 Jul 2026

PHREEQC-MCQ-200: A Diagnostic Benchmark for Tool-Augmented Scientific Simulator Agents

Model ReleasesDGX agent

arXiv:2607.00436v1 Announce Type: new Abstract: Large language model agents are increasingly connected to scientific software, yet it remains unclear when tool access makes scientific computation more

30 Jun 2026

Agentic Tool Use in Large Language Models

SafetyDGX agent

arXiv:2604.00835v2 Announce Type: replace Abstract: Large language models are increasingly being deployed as autonomous agents yet their real world effectiveness depends on reliable tools for informat

Dynamo: Dynamic Skill-Tool Evolution for Vision-Language Agents

Model ReleasesDGX agent

arXiv:2606.30185v1 Announce Type: new Abstract: Improving vision-language models (VLMs) on visual reasoning typically requires retraining or hand-designed prompts and tools. We present Dynamo, a train

8 Jun 2026

AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning

AgentsDGX agent

arXiv:2512.13278v2 Announce Type: replace Abstract: Agentic reinforcement learning has advanced large language models (LLMs) to reason through long chain-of-thought trajectories while interleaving ext

6 Jun 2026

good dev tools are cached intelligence for agents!

IndustryDGX agent

good dev tools are cached intelligence for agents! Token costs are why there will be no saas apocalypse / good dev tools are cached intelligence for agents! The popular theory goes: agents can write c

1 Jun 2026

Depth-Dependent Indirect Prompt Injection in Tool-Calling ReAct Agents: Injection Depth, Payload Framing, and Turn-Budget Sensitivity

Model ReleasesDGX agent

arXiv:2605.30686v1 Announce Type: cross Abstract: ReAct agents that interleave chain-of-thought reasoning with tool calls are increasingly deployed for real tasks such as scheduling, file retrieval, a

15 May 2026

From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents

Model ReleasesDGX agent

arXiv:2605.15104v1 Announce Type: new Abstract: Voice agents increasingly require reliable tool use from speech, whereas prominent tool-calling benchmarks remain text-based. We study whether verified

Latency-Quality Routing for Functionally Equivalent Tools in LLM Agents

Model ReleasesDGX agent

arXiv:2605.14241v1 Announce Type: new Abstract: Tool-augmented LLM agents increasingly access the same tool type through multiple functionally equivalent providers, such as web-search APIs, retrievers

13 May 2026

MCPShield: Content-Aware Attack Detection for LLM Agent Tool-Call Traffic

AgentsDGX agent

arXiv:2605.11053v1 Announce Type: cross Abstract: The Model Context Protocol (MCP) has become a widely adopted interface for LLM agents to invoke external tools, yet learned monitoring of MCP tool-cal

← Previous
123456…166
Next →