AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
5,141 results
Safety

ToolSelf: Unifying Task Execution and Self-Reconfiguration via Tool-Driven Emergent Adaptation

DGX agent

arXiv:2602.07883v3 Announce Type: replace Abstract: LLM-powered agentic systems excel at complex long-horizon tasks, but remain constrained by static configurations fixed before execution. Such rigidi

safetyarxiv-cs-ai
2 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

Evaluating using Mock Tool Calls to Quarantine Untrusted Prompt Inputs

DGX agent

arXiv:2605.30521v1 Announce Type: new Abstract: Large language models must frequently process untrusted inputs, such as judging an answer from another model or running tasks like spam and harm classif

researcharxiv-cs-cl
1 Jun 2026
Model Releases

VitalAgent: A Tool-Augmented Agent for Reactive and Proactive Physiological Monitoring over Wearable Health Data

DGX agent

arXiv:2605.29483v1 Announce Type: new Abstract: Wearable devices enable continuous monitoring of physiological signals such as ECG and PPG, but existing mHealth systems are largely limited to task-spe

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

MemCog: From Memory-as-Tool to Memory-as-Cognition in Conversational Agents

DGX agent

arXiv:2605.28046v1 Announce Type: new Abstract: Existing agent memory systems universally follow what we term a Memory-as-Tool paradigm where a single query triggers one-shot retrieval of flat passage

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

OralAgent: Integrating Reasoning, Tools, and Knowledge for Interactive Dental Image Analysis

DGX agent

arXiv:2605.27378v1 Announce Type: new Abstract: Dental image analysis plays a pivotal role in supporting accurate diagnosis and treatment planning in oral healthcare. Although recent advances have pro

model-releasesarxiv-cs-cl
28 May 2026
Safety

Agent-Facing Information Design in LLM Tool Registries

DGX agent

arXiv:2605.23916v1 Announce Type: cross Abstract: LLM tool registries function as unregulated advertising platforms: providers write free-text descriptions that agents use for selection, yet no measur

safetyarxiv-cs-ai
26 May 2026
Local Ai

DART: Semantic Recoverability for Structured Tool Agents

DGX agent

arXiv:2605.23311v1 Announce Type: new Abstract: When a structured tool agent fails mid-execution, the runtime faces a dilemma: replaying the entire task is safe but wasteful, while restoring from a lo

local-aiarxiv-cs-ai
25 May 2026
Model Releases

Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval

DGX agent

arXiv:2605.23826v1 Announce Type: cross Abstract: Keyframe selection is a direct way to provide verifiable visual evidence for long-video question answering (QA). Queries differ in what they require,

model-releasesarxiv-cs-cl
25 May 2026
Safety

GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation

DGX agent

arXiv:2605.21605v1 Announce Type: new Abstract: Open-ended image generation is no longer a simple prompt-to-image problem. High-quality generation often requires an agent to combine a model's internal

safetyarxiv-cs-cv
22 May 2026
Model Releases

LongVT: Incentivizing 'Thinking with Long Videos' via Native Tool Calling

DGX agent

arXiv:2511.20785v3 Announce Type: replace Abstract: Large multimodal models (LMMs) have shown great potential for video reasoning with textual Chain-of-Thought. However, they remain vulnerable to hall

model-releasesarxiv-cs-cv
22 May 2026
Safety

Progressive Autonomy as Preference Learning: A Formalization of Trust Calibration for Agentic Tool Use

DGX agent

arXiv:2605.19151v1 Announce Type: new Abstract: We formalize trust calibration for agentic tool use (deciding when an automated agent's proposed action may execute autonomously versus require human ap

safetyarxiv-cs-ai
20 May 2026
Model Releases

Toward Scalable Verifiable Reward: Proxy State-Based Evaluation for Multi-turn Tool-Calling LLM Agents

DGX agent

arXiv:2602.16246v3 Announce Type: replace Abstract: Interactive large language model (LLM) agents operating via multi-turn dialogue and multi-step tool calling are increasingly used in production. Ben

model-releasesarxiv-cs-ai
14 May 2026
Model Releases

ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox

DGX agent

arXiv:2605.10787v1 Announce Type: new Abstract: Current LLM agents are proficient at calling isolated APIs but struggle with the 'last mile' of commercial software automation. In real-world scenarios,

model-releasesarxiv-cs-ai
12 May 2026
Safety

Open Ontologies: Tool-Augmented Ontology Engineering with Stable Matching Alignment

DGX agent

arXiv:2605.09184v1 Announce Type: new Abstract: We present Open Ontologies, an open-source ontology engineering system implemented in Rust that integrates LLM-driven construction with formal OWL reaso

safetyarxiv-cs-ai
12 May 2026
Research

SearchSkill: Teaching LLMs to Use Search Tools with Evolving Skill Banks

DGX agent

arXiv:2605.09038v1 Announce Type: new Abstract: Teaching language models to use search tools is not only a question of whether they search, but also of whether they issue good queries. This is especia

researcharxiv-cs-ai
12 May 2026
Model Releases

ReasonSTL: Bridging Natural Language and Signal Temporal Logic via Tool-Augmented Process-Rewarded Learning

DGX agent

arXiv:2605.06483v2 Announce Type: replace Abstract: Signal Temporal Logic (STL) is an expressive formal language for specifying spatio-temporal requirements over real-valued, real-time signals. It has

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

ORPilot: A Production-Oriented Agentic LLM-for-OR Tool for Optimization Modeling

DGX agent

arXiv:2605.02728v1 Announce Type: new Abstract: This paper presents ORPilot, an open-source agentic AI system that translates real-world business problems into solver-ready optimization models. Unlike

model-releasesarxiv-cs-ai
6 May 2026
Safety

Tatemae: Detecting Alignment Faking via Tool Selection in LLMs

DGX agent

arXiv:2604.26511v1 Announce Type: cross Abstract: Alignment faking (AF) occurs when an LLM strategically complies with training objectives to avoid value modification, reverting to prior preferences o

safetyarxiv-cs-ai
30 Apr 2026
Safety

When Annotators Disagree, Topology Explains: Mapper, a Topological Tool for Exploring Text Embedding Geometry and Ambiguity

DGX agent

arXiv:2510.17548v2 Announce Type: replace Abstract: Language models are often evaluated with scalar metrics like accuracy, but such measures fail to capture how models internally represent ambiguity,

safetyarxiv-cs-cl
30 Apr 2026
Agents

Dr. RTL: Autonomous Agentic RTL Optimization through Tool-Grounded Self-Improvement

DGX agent

arXiv:2604.14989v2 Announce Type: replace Abstract: Recent advances in large language models (LLMs) have sparked growing interest in automatic RTL optimization for better performance, power, and area

agentsarxiv-cs-ai
28 Apr 2026
Safety

Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization

DGX agent

arXiv:2511.14846v2 Announce Type: replace-cross Abstract: Training Large Language Models (LLMs) for multi-turn Tool-Integrated Reasoning (TIR) - where models iteratively reason, generate code, and ver

safetyarxiv-cs-cl
21 Apr 2026
Model Releases

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench

DGX agent

arXiv:2604.16706v1 Announce Type: cross Abstract: Automated evaluation of tool-using large language model (LLM) agents is widely assumed to be reliable, but this assumption has rarely been validated a

model-releasesarxiv-cs-cl
21 Apr 2026
Safety

Training Language Models to Use Prolog as a Tool

DGX agent

arXiv:2512.07407v2 Announce Type: replace Abstract: Language models frequently produce plausible yet incorrect reasoning traces that are difficult to verify. We investigate fine-tuning models to use P

safetyarxiv-cs-cl
21 Apr 2026
Model Releases

Understanding Tool-Augmented Agents for Lean Formalization: A Factorial Analysis

DGX agent

arXiv:2604.16538v1 Announce Type: cross Abstract: Automatic translation of natural language mathematics into faithful Lean 4 code is hindered by the fundamental dissonance between informal set-theoret

model-releasesarxiv-cs-lg
21 Apr 2026
Agents

EnvScaler: Scaling Tool-Interactive Environments for LLM Agent via Programmatic Synthesis

DGX agent

arXiv:2601.05808v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are expected to be trained to act as agents in various real-world environments, but this process relies on rich a

agentsarxiv-cs-ai
20 Apr 2026
Model Releases

UI-Copilot: Advancing Long-Horizon GUI Automation via Tool-Integrated Policy Optimization

DGX agent

arXiv:2604.13822v1 Announce Type: new Abstract: MLLM-based GUI agents have demonstrated strong capabilities in complex user interface interaction tasks. However, long-horizon scenarios remain challeng

model-releasesarxiv-cs-lg
16 Apr 2026
Safety

SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents

DGX agent

arXiv:2604.07791v2 Announce Type: replace Abstract: Recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) have demonstrated significant potential in single-turn reasoning tasks. Wit

safetyarxiv-cs-ai
14 Apr 2026
Model Releases

EigentSearch-Q+: Enhancing Deep Research Agents with Structured Reasoning Tools

DGX agent

arXiv:2604.07927v2 Announce Type: replace Abstract: Deep research requires reasoning over web evidence to answer open-ended questions, and it is a core capability for AI agents. Yet many deep research

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios

DGX agent

arXiv:2604.06742v1 Announce Type: cross Abstract: Large Language Models (LLMs) are driving a shift towards intent-driven development, where agents build complete software from scratch. However, existi

model-releasesarxiv-cs-ai
10 Apr 2026
Research

Epistemic Transfer in AI-Assisted Verification: A Framework and Evaluation Protocol

DGX agent

arXiv:2608.08882v1 Announce Type: cross Abstract: AI tools that help people judge online claims are usually evaluated while the tool is present. This paper asks a different question: after using such

researcharxiv-cs-ai
11 Aug 2026
Model Releases

VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model with Structured Visual Reasoning and Native Tool Use

DGX agent

arXiv:2608.08477v1 Announce Type: new Abstract: We present VectraYX-Vision-1B, a sub-2B vision-language model (VLM) for Spanish/LATAM cybersecurity imagery, coupling a frozen SigLIP-so400m encoder to

model-releasesarxiv-cs-cl
11 Aug 2026
Local Ai

CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

DGX agent

arXiv:2608.03890v1 Announce Type: cross Abstract: A clinically useful chest X-ray system must go beyond fluent report generation: it should classify findings with tunable decision thresholds, localize

local-aiarxiv-cs-ai
5 Aug 2026
Tutorials

Preliminary Guidelines for Using and Evaluating GenAI Tools to Support Systematic Literature Reviews

DGX agent

arXiv:2607.24991v1 Announce Type: cross Abstract: Context: Generative AI (GenAI) and Large Language Models (LLMs) are increasingly used for academic tasks in software engineering and beyond, including

tutorialsarxiv-cs-ai
29 Jul 2026
Local Ai

MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning

DGX agent

arXiv:2607.24064v1 Announce Type: cross Abstract: Recent Vision-Language Models (VLMs) have achieved remarkable success in visual understanding, driven by the growing availability of high-quality imag

local-aiarxiv-cs-ai
28 Jul 2026
Agents

Improving Molecular Property Prediction in Small Language Models Using Graph-based Tools

DGX agent

arXiv:2607.13115v1 Announce Type: new Abstract: Small language models (SLMs) have shown promise for zero-shot molecular property prediction from SMILES strings, yet they often suffer from structural b

agentsarxiv-cs-ai
16 Jul 2026
Model Releases

Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents

DGX agent

arXiv:2607.07474v1 Announce Type: cross Abstract: Agentic red-teaming benchmarks report whether an injected agent was compromised as a single bit: the attack succeeded, or it did not. We argue that th

model-releasesarxiv-cs-ai
9 Jul 2026
Tutorials

Fast, Slow, and Tool-augmented Thinking for LLMs: A Review

DGX agent

arXiv:2508.12265v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated remarkable progress in reasoning across diverse domains. However, effective reasoning in real-world t

tutorialsarxiv-cs-cl
9 Jul 2026
Research

CSB: A Counting and Sampling tool for Bit-vectors

DGX agent

arXiv:2607.04142v1 Announce Type: cross Abstract: Satisfiability modulo theory (SMT) solvers have significantly advanced automated reasoning due to their effectiveness in solving problems across vario

researcharxiv-cs-ai
7 Jul 2026
Model Releases

Toward Efficient Agents: Memory, Tool learning, and Planning

DGX agent

arXiv:2601.14192v2 Announce Type: replace Abstract: Recent years have witnessed increasing interest in extending large language models into agentic systems. While the effectiveness of agents has conti

model-releasesarxiv-cs-ai
7 Jul 2026
Research

Why3-py: A Tool for Formal Verification of Hypothesis Testing and Meta-Analysis in Python

DGX agent

arXiv:2607.03951v1 Announce Type: cross Abstract: The reproducibility crisis in scientific research has received widespread recognition, thereby increasing the importance of meta-analyses that integra

researcharxiv-cs-ai
7 Jul 2026
Safety

Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows

DGX agent

arXiv:2607.01465v1 Announce Type: new Abstract: Large language models are trained to predict the next token, not to act inside a specific API. In niche enterprise SaaS workflows -- where success means

safetyarxiv-cs-ai
3 Jul 2026
Model Releases

Beyond Document Grounding: Span-Level Hallucination Detection over Code, Tool Output, and Documents

DGX agent

arXiv:2607.00895v1 Announce Type: new Abstract: Hallucination detection for retrieval-augmented generation (RAG) is usually evaluated on natural-language document evidence. However, grounded generatio

model-releasesarxiv-cs-cl
2 Jul 2026
Model Releases

Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use

DGX agent

arXiv:2607.01084v1 Announce Type: new Abstract: While Large Language Model (LLM) agents demonstrate proficiency in static benchmarks, their deployment in real-world scenarios is hindered by the dynami

model-releasesarxiv-cs-ai
2 Jul 2026
Research

An Integrated Two-Stage Deep-Learning Tool for Rapid Post-Hurricane Damage Identification and Repair Scheduling

DGX agent

arXiv:2606.29117v1 Announce Type: cross Abstract: Post-hurricane damage assessment and repair scheduling can require computationally intensive simulation and optimization. This paper presents an integ

researcharxiv-cs-lg
30 Jun 2026
Model Releases

CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents

DGX agent

arXiv:2511.02734v3 Announce Type: replace Abstract: Current evaluations of Large Language Model (LLM) agents primarily emphasize task completion, often overlooking resource efficiency and adaptability

model-releasesarxiv-cs-ai
30 Jun 2026
Safety

Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems

DGX agent

arXiv:2606.28425v1 Announce Type: cross Abstract: Increasingly autonomous agentic AI systems pose novel multi-agent risks, such as secret collusion via covert communication channels. The natural defen

safetyarxiv-cs-ai
30 Jun 2026
Model Releases

Towards Automating Scientific Review with Google's Paper Assistant Tool

DGX agent

arXiv:2606.28277v1 Announce Type: cross Abstract: Artificial intelligence is driving a revolution in scientific discovery, accelerating everything from hypothesis generation to mathematical theorem pr

model-releasesarxiv-cs-ai
29 Jun 2026
Model Releases

Confidence-Aware Tool Orchestration for Robust Video Understanding

DGX agent

arXiv:2606.26904v1 Announce Type: cross Abstract: Video reasoning language models implicitly assume that every input frame is equally reliable. This leads to what we term the Blind Trust Problem: unde

model-releasesarxiv-cs-ai
26 Jun 2026
← Previous
1…7891011…108
Next →