AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
5,141 results
Agents

FitText: Evolving Agent Tool Ecologies via Memetic Retrieval

DGX agent

arXiv:2605.02411v1 Announce Type: cross Abstract: A semantic gap separates how users describe tasks from how tools are documented. As API ecosystems scale to tens of thousands of endpoints, static ret

agentsarxiv-cs-lg
5 May 2026
Agents
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

ToolGrad: Efficient Tool-use Dataset Generation with Textual 'Gradients'

DGX agent

arXiv:2508.04086v2 Announce Type: replace Abstract: Prior work synthesizes tool-use LLM datasets by first generating a user query, followed by complex tool-use annotations like depth-first search (DFS

agentsarxiv-cs-cl
4 May 2026
Model Releases

Beyond Itinerary Planning-A Real-World Benchmark for Multi-Turn and Tool-Using Travel Tasks

DGX agent

arXiv:2512.22673v3 Announce Type: replace Abstract: Travel planning is a natural real-world task to test large language models' (LLMs) planning and tool-use abilities. Although prior work has studied

model-releasesarxiv-cs-ai
22 Apr 2026
Research

A Unified Compliance Aggregator Framework for Automated Multi-Tool Security Assessment of Linux Systems

DGX agent

arXiv:2604.17256v1 Announce Type: cross Abstract: Assessing the security posture of modern computing systems typically requires the use of multiple specialized tools. These tools focus on different as

researcharxiv-cs-lg
21 Apr 2026
Agents

OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning

DGX agent

arXiv:2502.11271v2 Announce Type: replace-cross Abstract: Solving complex reasoning tasks may involve visual understanding, domain knowledge retrieval, numerical calculation, and multi-step reasoning.

agentsarxiv-cs-cl
15 Apr 2026
Model Releases

The Amazing Agent Race: Strong Tool Users, Weak Navigators

DGX agent

arXiv:2604.10261v1 Announce Type: new Abstract: Existing tool-use benchmarks for LLM agents are overwhelmingly linear: our analysis of six benchmarks shows 55 to 100% of instances are simple chains of

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

UniToolCall: Unifying Tool-Use Representation, Data, and Evaluation for LLM Agents

DGX agent

arXiv:2604.11557v1 Announce Type: new Abstract: Tool-use capability is a fundamental component of LLM agents, enabling them to interact with external systems through structured function calls. However

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them

DGX agent

arXiv:2607.27703v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task

model-releasesarxiv-cs-ai
5 Aug 2026
Safety

TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning

DGX agent

arXiv:2608.04007v1 Announce Type: cross Abstract: Tool-Integrated Reasoning (TIR) enables LLMs to solve complex tasks through iterative tool interactions. However, existing reinforcement learning meth

safetyarxiv-cs-ai
5 Aug 2026
Agents

Scores Are Not Decisions: Cost-Aware Stopping for Tool Acquisition in LLM Agents

DGX agent

arXiv:2607.27083v1 Announce Type: new Abstract: As LLM agents increasingly depend on diverse external services such as search engines, databases, and connectors, agent harnesses face a fundamental too

agentsarxiv-cs-lg
30 Jul 2026
Agents

Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA Workflows

DGX agent

arXiv:2607.17528v3 Announce Type: replace Abstract: Large language model (LLM) agents are extending electronic design automation (EDA) beyond static RTL generation toward long-horizon, tool-interactiv

agentsarxiv-cs-ai
24 Jul 2026
Model Releases

A Systematic Evaluation of Traditional Privacy Policy Analysis Tools Against LLMs

DGX agent

arXiv:2607.17075v2 Announce Type: replace-cross Abstract: The advent of LLMs has significantly changed the research on privacy policy and data compliance analysis by enabling tasks that previously req

model-releasesarxiv-cs-cl
23 Jul 2026
Agents

CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration

DGX agent

arXiv:2607.05465v1 Announce Type: cross Abstract: Complex image creation and editing often require more than a single generation or editing model. A user request may involve synthesizing images, local

agentsarxiv-cs-ai
8 Jul 2026
Model Releases

FORGE: Towards Functional Tool-Use Generalization via Keypoint Trajectory Reasoning

DGX agent

arXiv:2607.05780v1 Announce Type: cross Abstract: While humans readily repurpose a book, a stone, or a shoe to drive a nail, robots trained on specific tools fail to transfer the same function to nove

model-releasesarxiv-cs-ai
8 Jul 2026
Safety

Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory

DGX agent

arXiv:2512.07287v3 Announce Type: replace-cross Abstract: As intents unfold and environments change, multi-turn agents face continuously shifting decision contexts. Although reusing past experience is

safetyarxiv-cs-ai
30 Jun 2026
Model Releases

Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models

DGX agent

arXiv:2601.05366v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly deployed as agents that invoke external tools through structured function calls. While recent wo

model-releasesarxiv-cs-ai
30 Jun 2026
Agents

Localizing RL-Induced Tool Use to a Single Crosscoder Feature

DGX agent

arXiv:2606.26474v1 Announce Type: cross Abstract: Fine-tuning through RL reshapes the internal representations of language models to enable agentic behaviors such as tool use, yet the mechanistic basi

agentsarxiv-cs-ai
26 Jun 2026
Safety

Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It

DGX agent

arXiv:2606.26027v1 Announce Type: new Abstract: Tool use enables large language models (LLMs) to perform complex tasks, and recent agentic reinforcement learning (RL) methods show promise for enhancin

safetyarxiv-cs-cl
25 Jun 2026
Model Releases

Self-Evolution for Multi-Turn Tool-Calling Agents via Divergence-Point Preference Learning

DGX agent

arXiv:2606.23112v1 Announce Type: new Abstract: Multi-turn tool-using agents must coordinate long-horizon tool sequences while tracking dialogue state and policy constraints. Existing approaches often

model-releasesarxiv-cs-lg
23 Jun 2026
Model Releases

ASA: Backbone-Training-Free Representation Engineering for Tool-Calling Agents

DGX agent

arXiv:2602.04935v3 Announce Type: replace-cross Abstract: Adapting LLM agents to domain-specific tool calling remains notably brittle under evolving interfaces. Prompt and schema engineering is easy t

model-releasesarxiv-cs-ai
10 Jun 2026
Agents

Exploring Agentic Tool-Calling Decisions via Uncertainty-Aligned Reinforcement Learning

DGX agent

arXiv:2606.06976v1 Announce Type: new Abstract: Large language model (LLM)-based agents often make suboptimal tool-use decisions, including unsupported tool invocation and hallucinated direct response

agentsarxiv-cs-ai
8 Jun 2026
Model Releases

TAPO: Tool-Aware Policy Optimization via Credit Transfer for Multimodal Search Agents

DGX agent

arXiv:2606.05784v1 Announce Type: new Abstract: We identify and formally characterize credit misassignment as a systematic failure mode of GRPO in tool-augmented multimodal search agents: its uniform

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

When Tools Fail: Benchmarking Dynamic Replanning and Anomaly Recovery in LLM Agents

DGX agent

arXiv:2606.05806v1 Announce Type: new Abstract: Existing benchmarks evaluate Tool-Integrated Reasoning (TIR) in LLMs on idealized ''happy paths'', largely overlooking real-world tool failures. We intr

model-releasesarxiv-cs-ai
6 Jun 2026
Research

DiG-Plan: Mitigating Early Commitment for Tool-Graph Planning via Diffusion Guidance

DGX agent

arXiv:2606.05728v1 Announce Type: cross Abstract: Generating executable tool plans requires selecting appropriate subsets from tool libraries, a combinatorial search problem with an exponentially larg

researcharxiv-cs-cl
5 Jun 2026
Agents

Atomix: Timely, Transactional Tool Use for Reliable Agentic Workflows

DGX agent

arXiv:2602.14849v2 Announce Type: replace-cross Abstract: LLM agents execute multi-step workflows that mutate external state through tools. Common orchestrators treat tool return as the settlement tri

agentsarxiv-cs-ai
2 Jun 2026
Safety

On Effectiveness and Efficiency of Agentic Tool-calling and RL Training

DGX agent

arXiv:2606.00135v1 Announce Type: cross Abstract: Tool-calling is a central component of modern large language model (LLM) agents, equipping them with skills beyond their parametric knowledge. This pa

safetyarxiv-cs-ai
2 Jun 2026
Research

DeepTool: Scaling Interleaved Deliberation in Tool-Integrated Reasoning via Process-Supervised Reinforcement Learning

DGX agent

arXiv:2605.29568v1 Announce Type: new Abstract: Tool-Integrated Reasoning (TIR) extends LLM capabilities by leveraging external environments. However, existing methods lack the deliberation during seq

researcharxiv-cs-ai
29 May 2026
Safety

Peak-Then-Collapse and the Four Interface Channels of Knowledge-Graph Tool Use

DGX agent

arXiv:2605.26037v1 Announce Type: new Abstract: We test the standard RLVR tool-use recipe -- GRPO on Qwen2.5-7B-Instruct -- on a deliberately minimal knowledge-graph tool API: four Freebase navigation

safetyarxiv-cs-cl
26 May 2026
Model Releases

ToolRegistry: A Protocol-Agnostic Tool Management Library for Function-Calling LLMs

DGX agent

arXiv:2507.10593v3 Announce Type: replace-cross Abstract: Every LLM tool call is structurally an RPC -- a function name, JSON arguments, and a serialized result -- yet each protocol (native Python, MC

model-releasesarxiv-cs-ai
26 May 2026
Applications

It's not the Language Model, it's the Tool: Deterministic Mediation for Scientific Workflows

DGX agent

arXiv:2605.13245v1 Announce Type: new Abstract: Language models can produce convincing scientific analyses, but repeated generations on the same data do not guarantee the same result. A researcher may

applicationsarxiv-cs-ai
14 May 2026
Model Releases

DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents

DGX agent

arXiv:2605.09679v1 Announce Type: cross Abstract: Medical vision-language models (VLMs) and AI agents have made significant progress in learning to analyze and reason about clinical images. However, e

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue Agent

DGX agent

arXiv:2601.18700v2 Announce Type: replace Abstract: Emotional Support Conversation requires not only affective expression but also grounded instrumental support to provide trustworthy guidance. Howeve

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Tools as Continuous Flow for Evolving Agentic Reasoning

DGX agent

arXiv:2605.07339v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in orchestrating tools for reasoning tasks. However, existing methods rely on a s

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use

DGX agent

arXiv:2605.02964v1 Announce Type: new Abstract: Reinforcement learning (RL) trained language model agents with tool access are increasingly deployed in coding assistants, research tools, and autonomou

model-releasesarxiv-cs-lg
6 May 2026
Model Releases

GAZE: Grounded Agentic Zero-shot Evaluation with Viewer-Level Tools and Literature Retrieval on Rare Brain MRI

DGX agent

arXiv:2605.00876v1 Announce Type: cross Abstract: Vision-language models (VLMs) read an image and produce text in a single forward pass, whereas radiologists typically inspect an image several times a

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

Reinforced Agent: Inference-Time Feedback for Tool-Calling Agents

DGX agent

arXiv:2604.27233v1 Announce Type: new Abstract: Tool-calling agents are evaluated on tool selection, parameter accuracy, and scope recognition, yet LLM trajectory assessments remain inherently post-ho

model-releasesarxiv-cs-ai
1 May 2026
Research

Tool Learning Needs Nothing More Than a Free 8B Language Model

DGX agent

arXiv:2604.17739v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a prevalent paradigm for training tool calling agents, which typically requires online interactive environments

researcharxiv-cs-cl
21 Apr 2026
Local Ai

Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments

DGX agent

arXiv:2508.08791v3 Announce Type: replace Abstract: Effective tool use is essential for large language models (LLMs) to interact with their environment. However, progress is limited by the lack of eff

local-aiarxiv-cs-cl
16 Apr 2026
Agents

Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models

DGX agent

arXiv:2604.08545v1 Announce Type: new Abstract: The advent of agentic multimodal models has empowered systems to actively interact with external environments. However, current agents suffer from a pro

agentsarxiv-cs-cv
10 Apr 2026
Agents

FlowScout: From Execution Feedback to Reliable Tool-Using Agent Workflows

DGX agent

arXiv:2608.10039v1 Announce Type: new Abstract: Agentic workflows have become an important abstraction for building reliable LLM-based automation systems by organizing large language models (LLMs), to

agentsarxiv-cs-lg
12 Aug 2026
Agents

MIRA: Medical Image Reflection for Agentic Diagnosis

DGX agent

arXiv:2608.10827v1 Announce Type: cross Abstract: Medical visual agents can use tools to inspect images and retrieve external knowledge, but indiscriminate tool use may introduce noisy or misleading e

agentsarxiv-cs-ai
12 Aug 2026
Safety

SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning

DGX agent

arXiv:2608.07959v1 Announce Type: new Abstract: Ultra-long egocentric video understanding requires reasoning over temporally sparse evidence distributed across hours or days, challenging current multi

safetyarxiv-cs-ai
11 Aug 2026
Agents

Self-Evolving Neuro-Symbolic Skills for Tool-Augmented Spatial Reasoning

DGX agent

arXiv:2608.07955v1 Announce Type: new Abstract: Large vision-language models have achieved strong performance in multimodal reasoning, but they remain unreliable on fine-grained spatial tasks that dem

agentsarxiv-cs-ai
11 Aug 2026
Research

Tools to Explain Neural Networks for Power System Dynamics

DGX agent

arXiv:2608.08048v1 Announce Type: cross Abstract: This paper presents, for the first time in power systems literature to our knowledge, analytical tools to explain the training performance of machine

researcharxiv-cs-ai
11 Aug 2026
Safety

IACM-RL: Intent-Aware Context Management and Reinforcement Learning for Complex Tool Invocation under Dynamic Intent Fluctuations

DGX agent

arXiv:2608.02110v1 Announce Type: new Abstract: Executing long-horizon tool invocations in real-world environments is severely challenged by dynamic user intent noise. Existing methods attempt robustn

safetyarxiv-cs-cl
4 Aug 2026
Model Releases

OoO-Spec: Out-of-Order Semantic Speculation for Fast Tool Calling

DGX agent

arXiv:2608.00814v1 Announce Type: new Abstract: LLMs generate tool calls token by token, even though the function choice and argument values can often be predicted in parallel from the request and too

model-releasesarxiv-cs-cl
4 Aug 2026
Agents

Beacon: Knowing When and How to Perform Agentic Visual Reasoning

DGX agent

arXiv:2607.28595v1 Announce Type: new Abstract: The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather tha

agentsarxiv-cs-cv
31 Jul 2026
Local Ai

In-the-Flow Agentic System Optimization for Effective Planning and Tool Use

DGX agent

arXiv:2510.05592v2 Announce Type: replace Abstract: Outcome-driven reinforcement learning has advanced reasoning in large language models (LLMs), but prevailing tool-augmented approaches train a singl

local-aiarxiv-cs-ai
23 Jul 2026
← Previous
1…34567…108
Next →