AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
5,141 results
Agents

OmniReasoner: Thinking with Long Audio-Video via Native Tool Use

DGX agent

arXiv:2607.19339v1 Announce Type: new Abstract: Long audio-video reasoning is difficult for omnimodal LLMs because the decisive evidence is often sparse, cross-modal, and too expensive to preserve wit

agentsarxiv-cs-cv
23 Jul 2026
Agents
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Early Adoption of Agentic Coding Tools by GitHub Projects

DGX agent

arXiv:2607.14037v1 Announce Type: cross Abstract: Agentic coding tools are increasingly capable of generating and submitting pull requests (PRs) to software projects, introducing new forms of human-ag

agentsarxiv-cs-ai
16 Jul 2026
Agents

From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents

DGX agent

arXiv:2607.07321v1 Announce Type: new Abstract: Tool utilization enables Large Language Model (LLM) agents to interact with the real world and resolve complex tasks. However, existing agent frameworks

agentsarxiv-cs-ai
9 Jul 2026
Agents

RIMRULE: Improving Tool-Using Language Agents via MDL-Guided Rule Learning

DGX agent

arXiv:2601.00086v3 Announce Type: replace Abstract: Large language models (LLMs) often struggle to use tools reliably in domain-specific settings, where APIs may be idiosyncratic, under-documented, or

agentsarxiv-cs-cl
9 Jul 2026
Model Releases

Beyond the Leaderboard: A Synthesis of Tool-Use, Planning, and Reasoning Failures in Large Language Model Agents

DGX agent

arXiv:2607.05775v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly evaluated on their ability to use tools, plan multi-step tasks, coordinate with other agents, and ope

model-releasesarxiv-cs-ai
8 Jul 2026
Agents

Unicode TAG-Block Concealment of Tool-Metadata Payloads in the Model Context Protocol: An Approval-View Fidelity Gap Across Three Independent Server Implementations

DGX agent

arXiv:2607.05744v1 Announce Type: cross Abstract: The Model Context Protocol (MCP) is the dominant way coding agents discover and invoke external tools. A server advertises each tool through a tools/l

agentsarxiv-cs-ai
8 Jul 2026
Model Releases

AgentLTL: A Trace-Verification Framework for Measuring, Enforcing, and Training Procedural Compliance in Tool-Using LLM Agents

DGX agent

arXiv:2607.02599v1 Announce Type: cross Abstract: Tool-using LLM agents are usually evaluated by final-answer correctness or LLM judges. Neither captures how an answer was produced. In safety-critical

model-releasesarxiv-cs-ai
7 Jul 2026
Local Ai

Reasoning effort, not tool access, buys first-try reliability in agentic code generation: an observational study

DGX agent

arXiv:2607.02436v1 Announce Type: cross Abstract: Agentic coding assistants are increasingly given extra capabilities, such as browser based testing tools and design oriented system prompts, on the as

local-aiarxiv-cs-ai
3 Jul 2026
Safety

Speech Playground: An Interactive Tool for Speech Analysis and Comparison

DGX agent

arXiv:2607.00418v1 Announce Type: new Abstract: This paper presents Speech Playground, an interactive speech visualization and comparison tool. While existing tools such as Praat are excellent, it can

safetyarxiv-cs-cl
2 Jul 2026
Model Releases

Artificial Intelligence in Sports: Insights from a Quantitative Survey among Sports Students in Germany about their Perceptions, Expectations, and Concerns regarding the Use of AI Tools

DGX agent

arXiv:2503.05785v2 Announce Type: replace-cross Abstract: Generative Artificial Intelligence (AI) tools such as ChatGPT, Copilot, or Gemini have a crucial impact on academic research and teaching. Emp

model-releasesarxiv-cs-ai
1 Jul 2026
Hardware

Optimizing CUDA like a Human: Micro-Profiling Tools as Expert Surrogates for LLM-Based GPU Kernel Optimization

DGX agent

arXiv:2606.26453v1 Announce Type: new Abstract: We present KernelPro, a closed-loop multi-agent system that automatically generates, profiles, and iteratively optimizes GPU kernel code by integrating

hardwarearxiv-cs-lg
26 Jun 2026
Model Releases

AgenticNav: Zero-Shot Vision-and-Language Navigation as a Tool-Calling Harness

DGX agent

arXiv:2606.10577v1 Announce Type: new Abstract: Zero-shot vision-and-language navigation in continuous environments (VLN-CE) has recently become feasible with large vision-language models (VLMs). Howe

model-releasesarxiv-cs-ro
10 Jun 2026
Safety

Context-Fractured Decomposition Attacks on Tool-Using LLM Agents: Exploiting Artifact Provenance Gaps

DGX agent

arXiv:2606.09084v1 Announce Type: cross Abstract: Tool-using LLM agents interact with the world through actions that persist state in artifacts (e.g., workspace files or logs). Consequently, jailbreak

safetyarxiv-cs-ai
9 Jun 2026
Agents

QueryWeaver: Reliable Multi-Tool Query Execution Planning via LLM-Based Graph Generation

DGX agent

arXiv:2606.08300v1 Announce Type: new Abstract: Many real-world queries over personal data span multiple applications and require structured planning, as individual tools expose only partial informati

agentsarxiv-cs-lg
9 Jun 2026
Model Releases

Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Use

DGX agent

arXiv:2603.03205v2 Announce Type: replace Abstract: Agentic language models operate in a fundamentally different safety regime than chat models: they must plan, call tools, and execute long-horizon ac

model-releasesarxiv-cs-cl
4 Jun 2026
Model Releases

AgentProcessBench: Diagnosing Step-Level Process Quality in Tool-Using Agents

DGX agent

arXiv:2603.14465v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have evolved into tool-using agents, they remain brittle in long-horizon interactions. Unlike mathematical reason

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

MAVEN: Improving Generalization in Agentic Tool Calling

DGX agent

arXiv:2605.30738v1 Announce Type: new Abstract: Generalization across agentic tool-calling environments remains a central challenge for reliable agentic reasoning systems. Although large language mode

model-releasesarxiv-cs-ai
1 Jun 2026
Agents

How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines

DGX agent

arXiv:2605.28840v1 Announce Type: cross Abstract: Large language model (LLM) agents with tool-calling capabilities are increasingly deployed in production systems, yet a fundamental reliability questi

agentsarxiv-cs-ai
29 May 2026
Model Releases

Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning

DGX agent

arXiv:2602.00994v2 Announce Type: replace Abstract: Agentic Reinforcement Learning (ARL) trains large language models to interleave reasoning with external tool execution to solve complex tasks. Most

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

DisasterBench: Benchmarking LLM Planning under Typed Tool Interface Constraints

DGX agent

arXiv:2605.27957v1 Announce Type: new Abstract: Disasters cause severe societal impacts, demanding rapid coordination of heterogeneous AI tools, from satellite analysis to flood prediction and damage

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

EgoBench: An Interactive Egocentric Multimodal Benchmark for Tool-Using Agents

DGX agent

arXiv:2605.27820v1 Announce Type: new Abstract: As AI agents increasingly operate in open, real-world environments, they require a deep synergy of multimodal perception, tool invocation with multi-hop

model-releasesarxiv-cs-ai
28 May 2026
Agents

Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use

DGX agent

arXiv:2605.27788v1 Announce Type: cross Abstract: Humans know when to reach for help e.g. 347 imes 28 warrants a calculator while 2+2 does not. Language models do not. Prompt-based approaches can inst

agentsarxiv-cs-cl
28 May 2026
Safety

Voluntary Collusion with Secret Tools in Competing LLM Agents

DGX agent

arXiv:2605.27593v1 Announce Type: new Abstract: Even when a tool is explicitly described as unfair and harmful to others, ostensibly safety-aligned LLM agents still voluntarily engage in secret collus

safetyarxiv-cs-ai
28 May 2026
Agents

Stateful Inference for Low-Latency Multi-Agent Tool Calling

DGX agent

arXiv:2605.26289v1 Announce Type: new Abstract: Multi-agent tool calling is becoming the dominant interaction pattern for LLM-based systems, yet existing inference frameworks treat each tool call as a

agentsarxiv-cs-lg
27 May 2026
Model Releases

AgroTools: A Benchmark for Tool-Augmented Multimodal Agents in Agriculture

DGX agent

arXiv:2605.22366v1 Announce Type: new Abstract: Agricultural decision-making increasingly requires multimodal systems that can transform visual observations into reliable, executable actions. However,

model-releasesarxiv-cs-cv
22 May 2026
Agents

SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations

DGX agent

arXiv:2605.22564v1 Announce Type: new Abstract: Today, tool-calling agents are commonly evaluated or tested on static datasets of execution traces, including input commands, agent responses, and assoc

agentsarxiv-cs-cl
22 May 2026
Agents

ChromaFlow: A Negative Ablation Study of Orchestration Overhead in Tool-Augmented Agent Evaluation

DGX agent

arXiv:2605.14102v1 Announce Type: new Abstract: Autonomous language-model agents increasingly combine planning, tool use, document processing, browsing, code execution, and verification loops. These c

agentsarxiv-cs-ai
15 May 2026
Safety

EGL-SCA: Structural Credit Assignment for Co-Evolving Instructions and Tools in Graph Reasoning Agents

DGX agent

arXiv:2605.10366v1 Announce Type: new Abstract: Graph reasoning agents operating from natural-language inputs must solve a coupled problem: they must reconstruct a structured graph instance from text,

safetyarxiv-cs-ai
12 May 2026
Model Releases

RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement

DGX agent

arXiv:2605.09730v1 Announce Type: new Abstract: Iterative self-refinement is a popular inference-time reliability technique, but its effectiveness in code-mode tool use depends heavily on the structur

model-releasesarxiv-cs-lg
12 May 2026
Agents

Switchcraft: AI Model Router for Agentic Tool Calling

DGX agent

arXiv:2605.07112v1 Announce Type: new Abstract: Agentic AI systems that invoke external tools are powerful but costly, leading developers to default to large models and overspend inference budgets. Mo

agentsarxiv-cs-ai
11 May 2026
Safety

Every Step Counts: Step-Level Credit Assignment for Tool-Integrated Text-to-SQL

DGX agent

arXiv:2605.04719v1 Announce Type: new Abstract: Tool-integrated Text-to-SQL parsing has emerged as a promising paradigm, framing SQL generation as a sequential decision-making process interleaved with

safetyarxiv-cs-cl
7 May 2026
Agents

Beyond State Machines: Executing Network Procedures with Agentic Tool-Calling Sequences

DGX agent

arXiv:2605.02584v1 Announce Type: cross Abstract: Agentic AI will be an essential enabling technology for designing future mobile communication systems, which could provide flexible and customized ser

agentsarxiv-cs-ai
6 May 2026
Safety

Semantic-Contact Fields for Category-Level Generalizable Tactile Tool Manipulation

DGX agent

arXiv:2602.13833v2 Announce Type: replace Abstract: Generalizing tool manipulation requires both semantic planning and precise physical control. Modern generalist robot policies, such as Vision-Langua

safetyarxiv-cs-ro
5 May 2026
Agents

TADI: Tool-Augmented Drilling Intelligence via Agentic LLM Orchestration over Heterogeneous Wellsite Data

DGX agent

arXiv:2605.00060v1 Announce Type: new Abstract: We present TADI (Tool-Augmented Drilling Intelligence), an agentic AI system that transforms drilling operational data into evidence-based analytical in

agentsarxiv-cs-ai
5 May 2026
Model Releases

When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors

DGX agent

arXiv:2604.21255v1 Announce Type: new Abstract: Model distillation is a primary driver behind the rapid progress of LLM agents, yet it often leads to behavioral homogenization. Many emerging agents sh

model-releasesarxiv-cs-cl
24 Apr 2026
Model Releases

SkillGraph: Graph Foundation Priors for LLM Agent Tool Sequence Recommendation

DGX agent

arXiv:2604.19793v1 Announce Type: new Abstract: LLM agents must select tools from large API libraries and order them correctly. Existing methods use semantic similarity for both retrieval and ordering

model-releasesarxiv-cs-ai
23 Apr 2026
Local Ai

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning

DGX agent

arXiv:2601.15724v2 Announce Type: replace Abstract: Long-form video understanding remains a fundamental challenge for current Video Large Language Models. Most existing models rely on static reasoning

local-aiarxiv-cs-cv
21 Apr 2026
Safety

Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception

DGX agent

arXiv:2510.23853v3 Announce Type: replace Abstract: Large language model (LLM) agents are increasingly used to interact with and execute tasks in dynamic environments. However, a critical yet overlook

safetyarxiv-cs-cl
17 Apr 2026
Model Releases

Long-Horizon Plan Execution in Large Tool Spaces through Entropy-Guided Branching

DGX agent

arXiv:2604.12126v1 Announce Type: new Abstract: Large Language Models (LLMs) have significantly advanced tool-augmented agents, enabling autonomous reasoning via API interactions. However, executing m

model-releasesarxiv-cs-ai
15 Apr 2026
Safety

The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment

DGX agent

arXiv:2604.12116v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as tool-augmented agents capable of executing system-level operations. While existing benchmarks

safetyarxiv-cs-ai
15 Apr 2026
Model Releases

The Scaffolding Matters More Than the Interface: A Controlled Comparison of MCP and CLI Tool Use Across Seven Agent Scaffoldings, Five Language Models, and One Software Task

DGX agent

arXiv:2608.08654v1 Announce Type: new Abstract: How much an AI coding agent costs to run can depend more on the agent scaffolding that drives it than on the interface through which it reaches its tool

model-releasesarxiv-cs-ai
11 Aug 2026
Safety

ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation

DGX agent

arXiv:2608.04436v1 Announce Type: new Abstract: Text-to-image (T2I) models can produce visually compelling images, yet they remain limited on open-world tasks that require complex semantic understandi

safetyarxiv-cs-cv
6 Aug 2026
Agents

Steganalysis of Adaptive Covert Collusion in Tool-Using Agent Populations: A Black-Box, Cross-Principal Approach

DGX agent

arXiv:2608.02698v1 Announce Type: cross Abstract: Tool-using agents built on large language models (LLMs) are increasingly deployed not by a single operator but by many, side by side on shared infrast

agentsarxiv-cs-ai
5 Aug 2026
Safety

DiffuseAgent-MI: Distributionally-Grounded,Tool-Integrated Self-Evolving Agents for Faithful Visual Reasoning

DGX agent

arXiv:2608.00540v1 Announce Type: new Abstract: Tool-integrated vision-language agents have made remarkable progress on compositional and multi-step visual reasoning. Yet their outputs frequently exhi

safetyarxiv-cs-cv
4 Aug 2026
Safety

Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales

DGX agent

arXiv:2607.25364v1 Announce Type: new Abstract: Tool-using agents expose structured calls but commonly attach free-form rationales. Such rationales are neither authorization nor reliable introspection

safetyarxiv-cs-ai
29 Jul 2026
Safety

Hybrid Analysis for Secure MCP Tool Use in LLM Agents

DGX agent

arXiv:2607.25297v1 Announce Type: cross Abstract: The rapid development of large language model (LLM) agents has enabled their broad adoption across diverse real-world tasks. To standardize interactio

safetyarxiv-cs-ai
29 Jul 2026
Agents

Toward Standardized Cross-Vendor Agent Tool Trust Management in Autonomous Networks

DGX agent

arXiv:2607.25914v1 Announce Type: new Abstract: Autonomous Network Levels 4-5 require AI agents to invoke tools across vendor boundaries without human oversight, yet existing management standards lack

agentsarxiv-cs-ai
29 Jul 2026
Model Releases

AlloBench: Measuring Online Tool Allocation Capability in LLM Agents

DGX agent

arXiv:2607.23332v1 Announce Type: new Abstract: Creating a reusable tool is an investment: an agent pays a fixed cost now in exchange for the potential of future reuse. Therefore, a user should prefer

model-releasesarxiv-cs-lg
28 Jul 2026
← Previous
1…45678…108
Next →