AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
5,141 results
Research

Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls

DGX agent

arXiv:2607.24343v1 Announce Type: cross Abstract: Language-model agents act through structured tool calls whose arguments carry different risks. Untrusted content may safely influence an email body bu

researcharxiv-cs-ai
28 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Agents

Ophiuchus: Incentivizing Tool-augmented 'Think with Images' for Joint Medical Segmentation, Understanding and Reasoning

DGX agent

arXiv:2512.14157v2 Announce Type: replace Abstract: Recent medical MLLMs have made significant progress in generating step-by-step textual reasoning chains. However, they still struggle with complex c

agentsarxiv-cs-ai
3 Jul 2026
Safety

Keep Policy Gradient in Charge: Sibling-Guided Credit Distillation for Long-Horizon Tool-Use Agents

DGX agent

arXiv:2606.12634v2 Announce Type: replace-cross Abstract: Long-horizon tool-use reinforcement learning learns from outcome verification, but trajectory-level advantages are broadcast over reasoning, A

safetyarxiv-cs-ai
1 Jul 2026
Model Releases

An AI agent for treatment reasoning over a biomedical tool universe

DGX agent

arXiv:2606.28692v1 Announce Type: new Abstract: Treatment reasoning underpins every therapeutic decision, integrating disease context, comorbidities, medications, contraindications, and evolving biome

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence

DGX agent

arXiv:2606.20515v2 Announce Type: replace Abstract: Real-world spatial intelligence requires reasoning over a continuous and evolving 3D world, yet existing VLMs and tool-augmented agents largely rema

model-releasesarxiv-cs-cv
30 Jun 2026
Safety

PairWise Image Finder: An Open-source Tool for Finding Visually Aligned Street-Level Image Pairs for Urban Perception Studies

DGX agent

arXiv:2606.08795v1 Announce Type: new Abstract: Change detection and scene recognition techniques have been widely applied to Street View Imagery (SVI) to understand changes in scenes across the years

safetyarxiv-cs-cv
9 Jun 2026
Model Releases

Adaptive Minds: Empowering Agents with LoRA-as-Tools

DGX agent

arXiv:2510.15416v2 Announce Type: replace Abstract: We investigate a framework in which LoRA adapters are treated as callable tools that a base language model can dynamically select and invoke. We hyp

model-releasesarxiv-cs-ai
4 Jun 2026
Safety

Formal Semantics for Agentic Tool Protocols: A Process Calculus Approach

DGX agent

arXiv:2603.24747v2 Announce Type: replace Abstract: The emergence of large language model agents capable of invoking external tools has created urgent need for formal verification of agent protocols.

safetyarxiv-cs-ai
4 Jun 2026
Agents

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents

DGX agent

arXiv:2602.12984v2 Announce Type: replace Abstract: Scientific reasoning inherently demands integrating sophisticated toolkits to navigate domain-specific knowledge. Yet, current benchmarks largely ov

agentsarxiv-cs-cl
2 Jun 2026
Agents

SkillSmith: Co-Evolving Skills and Tools for Self-Improving Agent Systems

DGX agent

arXiv:2606.01314v1 Announce Type: new Abstract: Recent self-evolving agents have shown that skills can be discovered, refined, and accumulated through execution. However, existing skill-evolution fram

agentsarxiv-cs-ai
2 Jun 2026
Model Releases

T1: Tool-integrated Verification for Test-time Compute Scaling in Small Language Models

DGX agent

arXiv:2504.04718v2 Announce Type: replace-cross Abstract: Recent studies have demonstrated that test-time compute scaling effectively improves the performance of small language models (sLMs). However,

model-releasesarxiv-cs-ai
2 Jun 2026
Research

CoHyDE: Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool Retrieval

DGX agent

arXiv:2605.29271v1 Announce Type: new Abstract: Tool retrieval over large API catalogs is a core bottleneck for LLM agents: user queries arrive in colloquial, often underspecified language, while the

researcharxiv-cs-ai
29 May 2026
Agents

Evo-Attacker: Memory-Augmented Reinforcement Learning for Long-Horizon Tool Attacks on LLM-MAS

DGX agent

arXiv:2605.25389v1 Announce Type: cross Abstract: While Large Language Model-based Multi-Agent Systems (LLM-MAS) demonstrate remarkable capabilities in solving complex tasks by orchestrating specializ

agentsarxiv-cs-ai
26 May 2026
Safety

B-GRTO: Bootstrapped Group Relative Tool Optimization for Referring Segmentation

DGX agent

arXiv:2605.23500v1 Announce Type: new Abstract: Segmentation is a fundamental task in computer vision, underpinning pixel-level scene understanding and serving as a cornerstone for applications rangin

safetyarxiv-cs-cv
25 May 2026
Model Releases

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs

DGX agent

arXiv:2605.19528v1 Announce Type: new Abstract: 3D localization in Multimodal Large Language Models (MLLMs), including 3D object detection and 3D visual grounding, is fundamentally limited by camera i

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

OmniVL-Guard Pro: A Tool-Augmented Agent for Omnibus Vision-Language Forensics

DGX agent

arXiv:2605.16962v1 Announce Type: cross Abstract: Existing vision-language forgery detection and grounding methods operate under a closed-world paradigm, assuming verification can be completed by the

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

TIER: Trajectory-Invariant Execution Rewards for Multi-Step Tool Composition

DGX agent

arXiv:2605.16790v1 Announce Type: cross Abstract: Tool use enables large language models to solve complex tasks through sequences of API calls, yet existing reinforcement learning approaches fail to s

model-releasesarxiv-cs-ai
19 May 2026
Agents

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents

DGX agent

arXiv:2510.02837v2 Announce Type: replace Abstract: Although recent tool-augmented benchmarks involve complex requests, evaluation remains limited to answer matching, neglecting critical trajectory as

agentsarxiv-cs-ai
15 May 2026
Agents

Case-Based Calibration of Adaptive Reasoning and Execution for LLM Tool Use

DGX agent

arXiv:2605.15041v1 Announce Type: new Abstract: Tool use extends large language models beyond parametric knowledge, but reliable execution requires balancing appropriate reasoning depth with strict st

agentsarxiv-cs-ai
15 May 2026
Model Releases

VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity Language Model with Curriculum Learning and Native Tool Use

DGX agent

arXiv:2605.13989v1 Announce Type: new Abstract: We present VectraYX-Nano, a 41.95M-parameter decoder-only language model trained from scratch in Spanish for cybersecurity, with a Latin-American focus

model-releasesarxiv-cs-cl
15 May 2026
Agents

AgentShield: Deception-based Compromise Detection for Tool-using LLM Agents

DGX agent

arXiv:2605.11026v1 Announce Type: cross Abstract: Defenses against indirect prompt injection (IPI) in tool-using LLM agents share two structural weaknesses. First, they all attempt to prevent attacks

agentsarxiv-cs-cl
13 May 2026
Research

Read, Extract, Classify: A Tool for Smarter Requirements Engineering

DGX agent

arXiv:2605.11045v1 Announce Type: cross Abstract: This paper presents the ReXCL tool, which automates the extraction and classification processes in requirements engineering, enhancing the software de

researcharxiv-cs-lg
13 May 2026
Tutorials

Do Agents Need to Plan Step-by-Step? Rethinking Planning Horizon in Data-Centric Tool Calling

DGX agent

arXiv:2605.08477v1 Announce Type: new Abstract: Explicit planning is a critical capability for LLM-based agents solving complex data-centric tasks, which require precise tool calling over external dat

tutorialsarxiv-cs-cl
12 May 2026
Agents

Security Risks in Tool-Enabled AI Agents: A Systematic Analysis of Privileged Execution Environments

DGX agent

arXiv:2605.09721v1 Announce Type: cross Abstract: Tool-enabled AI agents are increasingly deployed in cloud-hosted environments and offered as services, where they perform side-effecting operations th

agentsarxiv-cs-ai
12 May 2026
Model Releases

Stable Agentic Control: Tool-Mediated LLM Architecture for Autonomous Cyber Defense

DGX agent

arXiv:2605.03034v1 Announce Type: new Abstract: Agentic systems involved in high-stake decision-making under adversarial pressure need formal guarantees not offered by existing approaches. Motivated b

model-releasesarxiv-cs-ai
7 May 2026
Research

Mechanistic Interpretability Tool for AI Weather Models

DGX agent

arXiv:2604.20467v1 Announce Type: cross Abstract: Artificial Intelligence (AI) weather models are improving rapidly, and their forecasts are already competitive with long-established traditional Numer

researcharxiv-cs-lg
23 Apr 2026
Model Releases

ComPASS: Towards Personalized Agentic Social Support via Tool-Augmented Companionship

DGX agent

arXiv:2604.18356v1 Announce Type: new Abstract: Developing compassionate interactive systems requires agents to not only understand user emotions but also provide diverse, substantive support. While r

model-releasesarxiv-cs-cl
21 Apr 2026
Safety

Waking Up Blind: Cold-Start Optimization of Supervision-Free Agentic Trajectories for Grounded Visual Perception

DGX agent

arXiv:2604.17475v1 Announce Type: cross Abstract: Small Vision-Language Models (SVLMs) are efficient task controllers but often suffer from visual brittleness and poor tool orchestration. They typical

safetyarxiv-cs-cl
21 Apr 2026
Model Releases

AISysRev -- LLM-based Tool for Title-abstract Screening

DGX agent

arXiv:2510.06708v3 Announce Type: replace-cross Abstract: Conducting systematic reviews is laborious. In the screening or study selection phase, the number of papers can be overwhelming. Recent resear

model-releasesarxiv-cs-ai
20 Apr 2026
Model Releases

LAST: Leveraging Tools as Hints to Enhance Spatial Reasoning for Multimodal Large Language Models

DGX agent

arXiv:2604.09712v1 Announce Type: cross Abstract: Spatial reasoning is a cornerstone capability for intelligent systems to perceive and interact with the physical world. However, multimodal large lang

model-releasesarxiv-cs-ai
14 Apr 2026
Agents

AnomalyAgent: Agentic Industrial Anomaly Synthesis via Tool-Augmented Reinforcement Learning

DGX agent

arXiv:2604.07900v1 Announce Type: new Abstract: Industrial anomaly generation is a crucial method for alleviating the data scarcity problem in anomaly detection tasks. Most existing anomaly synthesis

agentsarxiv-cs-cv
10 Apr 2026
Agents

TOOLCAD: Exploring Tool-Using Large Language Models in Text-to-CAD Generation with Reinforcement Learning

DGX agent

arXiv:2604.07960v1 Announce Type: cross Abstract: Computer-Aided Design (CAD) is an expert-level task that relies on long-horizon reasoning and coherent modeling actions. Large Language Models (LLMs)

agentsarxiv-cs-cl
10 Apr 2026
Safety

Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents

DGX agent

arXiv:2608.11110v1 Announce Type: new Abstract: When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compa

safetyarxiv-cs-cl
12 Aug 2026
Safety

The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images

DGX agent

arXiv:2608.06270v1 Announce Type: new Abstract: The 'thinking-with-images' paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom. However, models using these operations o

safetyarxiv-cs-ai
7 Aug 2026
Model Releases

Speed Reading Tool Powered by Artificial Intelligence for Students with ADHD, Dyslexia, and Short Attention Span

DGX agent

arXiv:2307.14544v2 Announce Type: replace-cross Abstract: This paper presents an artificial intelligence tool designed to assist students with dyslexia, ADHD, and short attention spans in processing t

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use

DGX agent

arXiv:2607.20536v1 Announce Type: new Abstract: Tool-use agents that address day-to-day digital tasks such as ordering groceries must not only operate applications, but also interact with the user, e.

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Code-in-the-Loop Forensics: Agentic Tool Use for Image Forgery Detection

DGX agent

arXiv:2512.16300v3 Announce Type: replace Abstract: Existing image forgery detection (IFD) methods either exploit low-level, semantics-agnostic artifacts or rely on multimodal large language models (M

model-releasesarxiv-cs-ai
23 Jul 2026
Agents

Personalized Recommendation Tool Learning via Autonomous Language Agents

DGX agent

arXiv:2607.19739v1 Announce Type: cross Abstract: Although large language models (LLMs) have recently gained traction in recommender systems due to their strong reasoning capabilities and extensive wo

agentsarxiv-cs-ai
23 Jul 2026
Model Releases

Compaction as Epistemic Failure: How Agentic LLM Tools Fabricate Confirmed Results from Killed Processes

DGX agent

arXiv:2607.13071v1 Announce Type: cross Abstract: Agentic LLM coding tools compress long session histories into compaction summaries that subsequent sessions inherit as ground truth. This paper docume

model-releasesarxiv-cs-ai
16 Jul 2026
Research

IQA-T1: Tool-based Visual Evidence Reasoning for Image Quality Assessment

DGX agent

arXiv:2607.12375v1 Announce Type: cross Abstract: Image Quality Assessment (IQA) in open-world environments remains challenging due to limited generalization and interpretability. Recent approaches ba

researcharxiv-cs-ai
15 Jul 2026
Research

AI tools in Arab University English classrooms: Looking back and forward

DGX agent

arXiv:2607.05403v1 Announce Type: cross Abstract: This paper aims to synthesize empirical research on AI tools used to support English as a second/foreign language (EL2) learners in Arab University cl

researcharxiv-cs-ai
8 Jul 2026
Model Releases

VERITAS: Towards a General-Purpose Replication Tool for Scientific Research

DGX agent

arXiv:2607.02931v1 Announce Type: new Abstract: AI tools are accelerating scientific publication while the systems that review it struggle to keep up, and independent verification of published researc

model-releasesarxiv-cs-ai
7 Jul 2026
Research

Supporting Workflow Reproducibility by Linking Bioinformatics Tools across Papers and Executable Code

DGX agent

arXiv:2603.08195v2 Announce Type: replace Abstract: Motivation: The rapid growth of biological data has intensified the need for transparent, reproducible, and well-documented computational workflows.

researcharxiv-cs-cl
30 Jun 2026
Research

To Tab or Not to Tab: Measuring Critical Engagement in AI Code Completion Tools Using Behavioral Signals and Attention Checks

DGX agent

arXiv:2606.30549v1 Announce Type: cross Abstract: AI code completion tools, such as Github Copilot, provide students with code suggestions to help them write programs. However, recent qualitative stud

researcharxiv-cs-ai
30 Jun 2026
Agents

Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task

DGX agent

arXiv:2512.10359v1 Announce Type: cross Abstract: Video Question Answering (VideoQA) task serves as a critical playground for evaluating whether foundation models can effectively perceive, understand,

agentsarxiv-cs-ai
30 Jun 2026
Local Ai

Can Open-Source LLM Agents Replace Static Application Security Testing Tools? An Empirical Assessment

DGX agent

arXiv:2606.11672v1 Announce Type: cross Abstract: This paper explores the value of agentic AI tools for cybersecurity purposes. We evaluate the efficacy of a general-purpose GenAI Large Language Model

local-aiarxiv-cs-ai
11 Jun 2026
Safety

Translate-R1: Cost-Aware Translation Tool Use via Reinforcement Learning

DGX agent

arXiv:2606.06835v1 Announce Type: new Abstract: The performance gap across languages in LLMs is well documented, and closing it natively requires pretraining or fine-tuning on corpora that, for most l

safetyarxiv-cs-cl
8 Jun 2026
Research

SCI-PRM: A Tool Aware Process Reward Model for Scientific Reasoning Verification

DGX agent

arXiv:2606.04579v1 Announce Type: new Abstract: While Process Reward Models (PRMs) have achieved remarkable success in mathematical reasoning, their application in complex scientific domains-such as b

researcharxiv-cs-ai
4 Jun 2026
← Previous
1…56789…108
Next →