AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
5,141 results
Safety

ASAP: Agent-System Co-Design for Wall-Clock-Centered Auto HPO Research for ML Experiments

DGX agent

arXiv:2606.25207v1 Announce Type: cross Abstract: Hyperparameter Optimization (HPO) is essential for maximizing machine learning model performance, and its core challenge is sample efficiency: finding

safetyarxiv-cs-cl
25 Jun 2026
Agents
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Bayesian control for coding agents

DGX agent

arXiv:2606.24453v1 Announce Type: new Abstract: Modern coding agents pair LLM generators with various tools, including cheap diagnostics and expensive verifiers. The tool-use decisions are typically g

agentsarxiv-cs-ai
24 Jun 2026
Agents

PRInTS: Reward Modeling for Long-Horizon Information Seeking

DGX agent

arXiv:2511.19314v2 Announce Type: replace Abstract: Information-seeking is a core capability for AI agents, requiring them to gather and reason over tool-generated information across long trajectories

agentsarxiv-cs-ai
11 Jun 2026
Local Ai

AGENTSERVESIM: A Hardware-aware Simulator for Multi-Turn LLM Agent Serving

DGX agent

arXiv:2606.09613v1 Announce Type: cross Abstract: Multi-turn LLM agents interleave model calls with external tool invocations, shifting serving from stateless request processing to stateful program ex

local-aiarxiv-cs-ai
9 Jun 2026
Safety

From Agent Traces to Trust: Evidence Tracing and Execution Provenance in LLM Agents

DGX agent

arXiv:2606.04990v1 Announce Type: cross Abstract: Large language model (LLM)-based agents increasingly solve complex tasks by interacting with external tools, retrieval systems, memory modules, enviro

safetyarxiv-cs-ai
4 Jun 2026
Local Ai

Agent libOS: A Library-OS-Inspired Runtime for Long-Running, Capability-Controlled LLM Agents

DGX agent

arXiv:2606.03895v1 Announce Type: cross Abstract: Large language model (LLM) agents are evolving from request-response assistants into long-running software actors: they maintain state across model ca

local-aiarxiv-cs-ai
3 Jun 2026
Safety

ToolFG: Towards Well-Grounded Fine-Grained Image Classification

DGX agent

arXiv:2606.02518v1 Announce Type: new Abstract: Fine-grained image classification (FGIC) has broad applications and has attracted significant research attention. In this paper, we explore a novel para

safetyarxiv-cs-cv
2 Jun 2026
Model Releases

Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systems

DGX agent

arXiv:2605.29676v1 Announce Type: new Abstract: Large language models in Agentic AI systems consume tool schemas and execution results and emit tool invocations as structured data. The default languag

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks

DGX agent

arXiv:2605.28556v1 Announce Type: new Abstract: As agent capabilities advance, existing benchmarks, such as au^2-Bench, are becoming increasingly saturated. Yet constructing new benchmark tasks remain

model-releasesarxiv-cs-ai
28 May 2026
Safety

When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?

DGX agent

arXiv:2605.27932v1 Announce Type: cross Abstract: Think-with-image reasoning is emerging as a new inference paradigm for large vision-language models, but its safety implications remain poorly underst

safetyarxiv-cs-ai
28 May 2026
Model Releases

VisualNeedle: Benchmarking Active Visual Search in Information-Dense Scenes

DGX agent

arXiv:2605.26380v1 Announce Type: cross Abstract: Frontier multimodal large language models (MLLMs) have been reported to achieve over 90% accuracy on fine-grained perception benchmarks. However, such

model-releasesarxiv-cs-ai
27 May 2026
Agents

MCPXKIT: The Unified Toolkit for Analyzing Model Context Protocol Security

DGX agent

arXiv:2508.12538v2 Announce Type: replace-cross Abstract: The Model Context Protocol (MCP) has emerged as a universal standard that enables AI agents to seamlessly connect with external tools, signifi

agentsarxiv-cs-ai
26 May 2026
Agents

Towards Multi-Turn Dialog Systems for Industrial Asset Operations and Maintenance

DGX agent

arXiv:2605.24953v1 Announce Type: new Abstract: Industrial asset operations and maintenance question answering is inherently multi-turn, iterative, and highly dependent on external tool invocation. Ho

agentsarxiv-cs-ai
26 May 2026
Model Releases

CIVeX: Causal Intervention Verification for Language Agents

DGX agent

arXiv:2605.09168v1 Announce Type: new Abstract: A valid tool call is not necessarily a valid intervention. Tool-using language agents are guarded by schema validators, policy filters, provenance check

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents

DGX agent

arXiv:2605.10832v1 Announce Type: new Abstract: Multimodal deep search requires an agent to solve open-world problems by chaining search, tool use, and visual reasoning over evolving textual and visua

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning

DGX agent

arXiv:2605.07251v1 Announce Type: new Abstract: Large Language Models (LLMs) have become increasingly capable as tool-using agents, with benchmarks spanning diverse general agentic tasks. Yet rigorous

model-releasesarxiv-cs-ai
11 May 2026
Safety

SOD: Step-wise On-policy Distillation for Small Language Model Agents

DGX agent

arXiv:2605.07725v1 Announce Type: cross Abstract: Tool-integrated reasoning (TIR) is difficult to scale to small language models due to instability in long-horizon tool interactions and limited model

safetyarxiv-cs-ai
11 May 2026
Agents

Agentic AI for Remote Sensing: Technical Challenges and Research Directions

DGX agent

arXiv:2604.24919v1 Announce Type: new Abstract: Earth Observation (EO) is moving beyond static prediction toward multi-step analytical workflows that require coordinated reasoning over data, tools, an

agentsarxiv-cs-cv
29 Apr 2026
Safety

AgentBound: Securing Execution Boundaries of AI Agents

DGX agent

arXiv:2510.21236v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have evolved into AI agents that interact with external tools and environments to perform complex tasks. The Mode

safetyarxiv-cs-ai
27 Apr 2026
Safety

AgentGL: Towards Agentic Graph Learning with LLMs via Reinforcement Learning

DGX agent

arXiv:2604.05846v2 Announce Type: replace Abstract: Large Language Models (LLMs) increasingly rely on agentic capabilities-iterative retrieval, tool use, and decision-making-to overcome the limits of

safetyarxiv-cs-cl
24 Apr 2026
Model Releases

MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling

DGX agent

arXiv:2511.11793v3 Announce Type: replace Abstract: We present MiroThinker v1.0, an open-source research agent designed to advance tool-augmented reasoning and information-seeking capabilities. Unlike

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification

DGX agent

arXiv:2604.01687v2 Announce Type: replace Abstract: Anthropic proposes the concept of skills for LLM agents to tackle multi-step professional tasks that simple tool invocations cannot address. A tool

model-releasesarxiv-cs-ai
14 Apr 2026
Local Ai

Inside the Scaffold: A Source-Code Taxonomy of Coding Agent Architectures

DGX agent

arXiv:2604.03515v2 Announce Type: replace-cross Abstract: LLM-based coding agents can localize bugs, generate patches, and run tests with diminishing human oversight, yet the scaffolding code that sur

local-aiarxiv-cs-ai
14 Apr 2026
Model Releases

Structured Uncertainty guided Clarification for LLM Agents

DGX agent

arXiv:2511.08798v2 Announce Type: replace-cross Abstract: LLM agents with tool-calling capabilities often fail when user instructions are ambiguous or incomplete, leading to incorrect invocations and

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

Learning to Triage Vulnerability Reports from Program Analysis: An Empirical Study in Node.js

DGX agent

arXiv:2510.20739v2 Announce Type: replace-cross Abstract: Program analysis tools often produce large volumes of candidate vulnerability reports that require costly manual review, creating a practical

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

One Adapter Pair per Model: A Universal Activation Interface for Language Models

DGX agent

arXiv:2608.09521v1 Announce Type: new Abstract: Activation-based tools are usually tied to one model's native hidden space, requiring probes, sparse autoencoders, and natural-language interpreters to

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

You Don't Need To Stay in The Loop: An Agentic Robotics Loop for Robot-Policy Improvement

DGX agent

arXiv:2608.07555v1 Announce Type: new Abstract: Coding agents such as Claude Code and Codex close the software loop: a main agent manages the loop, subagents analyze and execute, tools do the work. We

model-releasesarxiv-cs-ro
11 Aug 2026
Agents

Agentic Planning for Symbolic Execution

DGX agent

arXiv:2608.06397v1 Announce Type: cross Abstract: Symbolic execution seeks to explore feasible program paths, yet a practical run may exhaust its resources while much program behaviour remains unreach

agentsarxiv-cs-ai
10 Aug 2026
Model Releases

WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks

DGX agent

arXiv:2608.03499v1 Announce Type: new Abstract: Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment targets: each user can be served b

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

SERL-SQL: Selective Hindsight Distillation for Text-to-SQL Reinforcement Agentic Learning

DGX agent

arXiv:2608.00485v1 Announce Type: new Abstract: Recent Text-to-SQL systems increasingly rely on multi-turn interaction, execution feedback, and reinforcement learning. However, most existing methods u

model-releasesarxiv-cs-cl
4 Aug 2026
Agents

SIEVE: Selective Integrity Verification and Escalation for Defending LLM Agents against Indirect Prompt Injection

DGX agent

arXiv:2512.06716v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly used as the core of agentic systems due to their strong reasoning, planning, and tool-use capabi

agentsarxiv-cs-cl
4 Aug 2026
Safety

MemTX: Transactional Belief Commit for Stateful Agent Memory

DGX agent

arXiv:2607.23929v1 Announce Type: new Abstract: LLM agents increasingly coordinate through persistent shared memory: one agent's write becomes another agent's premise, and eventually a tool call with

safetyarxiv-cs-ai
28 Jul 2026
Model Releases

SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows

DGX agent

arXiv:2607.23123v1 Announce Type: new Abstract: Existing evaluations of large language models cover knowledge, reasoning, coding, and tool use, but they rarely treat a verifiable deliverable produced

model-releasesarxiv-cs-ai
28 Jul 2026
Agents

Causal-AgentIR: Self-Evolving Causal Memory for Adaptive Image Restoration Agents

DGX agent

arXiv:2607.21125v1 Announce Type: new Abstract: Image restoration agents have recently emerged as a flexible paradigm for handling diverse and unpredictable degradations in real-world scenarios. Exist

agentsarxiv-cs-cv
24 Jul 2026
Model Releases

GuardianAgentBench: Where Agents Fail and How to Guard Them

DGX agent

arXiv:2607.20982v1 Announce Type: new Abstract: As large language model agents increasingly operate autonomously with access to tools and external environments, ensuring their safe and reliable behavi

model-releasesarxiv-cs-ai
24 Jul 2026
Agents

ChainWatch: A Kill Chain-Aligned Sequential Detection Framework for Multi-Step Attacks in MCP-Based AI Agent Systems

DGX agent

arXiv:2607.19432v1 Announce Type: cross Abstract: The Model Context Protocol (MCP) is an open-source standard that allows AI agents to connect to external tools, databases, and services. While this co

agentsarxiv-cs-ai
23 Jul 2026
Model Releases

First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations

DGX agent

arXiv:2512.01241v4 Announce Type: replace-cross Abstract: Large language models (LLMs) and medical AI tools are routinely used by physicians and patients for medical advice, yet their clinical safety

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

How Inference Compute Shapes Frontier LLM Evaluation

DGX agent

arXiv:2606.17930v2 Announce Type: replace Abstract: AI evaluations are shifting toward harder tasks that benefit from longer trajectories involving tool use and iterative problem solving. As a result,

model-releasesarxiv-cs-ai
15 Jul 2026
Local Ai

PalmClaw: A Native On-Device Agent Framework for Mobile Phones

DGX agent

arXiv:2607.13027v1 Announce Type: cross Abstract: Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling tools, observing the results, and it

local-aiarxiv-cs-ai
15 Jul 2026
Model Releases

GLM-5 Serving Parameter Tuning for OpenClaw: Single-Deployment MaaS Inference Optimization for Long-Context Agent Workloads

DGX agent

arXiv:2607.02518v1 Announce Type: cross Abstract: OpenClaw requests are dominated by long, tool-augmented prefixes, including system prompts, conversation history, and tool outputs fed back into the c

model-releasesarxiv-cs-ai
7 Jul 2026
Agents

Rethinking Scientific Discovery in an Agentic Era

DGX agent

arXiv:2607.03863v1 Announce Type: new Abstract: Artificial intelligence has advanced scientific discovery, but most AI4Science systems remain fragmented tools that rely on humans to coordinate problem

agentsarxiv-cs-cl
7 Jul 2026
Hardware

SPORK: Self-Speculative Forking to Accelerate Agentic LLM Inference

DGX agent

arXiv:2607.03333v1 Announce Type: cross Abstract: LLM agents are becoming a common interface for research, coding, and question answering, yet their Thought-Action-Observation loop is often serial: th

hardwarearxiv-cs-ai
7 Jul 2026
Safety

ElephantAgent: Contextual State Continuity in Agentic Systems

DGX agent

arXiv:2607.01919v1 Announce Type: new Abstract: Agentic systems enhance their capabilities by invoking external tools and maintaining persistent memory. However, these external dependencies introduce

safetyarxiv-cs-ai
3 Jul 2026
Safety

Safeguarding LLM Agents from Misalignment through Provenance Analysis

DGX agent

arXiv:2607.01236v1 Announce Type: cross Abstract: As LLM agents gain increasing access to powerful tools, ensuring that their actions are aligned with the user's intent becomes critical. When an agent

safetyarxiv-cs-ai
3 Jul 2026
Agents

Agentic-Ideation: Sample Efficient Agentic Trajectories Synthesis for Scientific Ideation Agents

DGX agent

arXiv:2606.31229v1 Announce Type: new Abstract: Ideation plays a pivotal role in scientific discovery. Recent LLM, especially AI Scientist systems, show promising potential for automated ideation. How

agentsarxiv-cs-ai
1 Jul 2026
Model Releases

MCP Server Architecture Patterns for LLM-Integrated Applications

DGX agent

arXiv:2606.30317v1 Announce Type: cross Abstract: The Model Context Protocol (MCP), introduced by Anthropic in November 2024, defines a standardized interface for connecting large language models (LLM

model-releasesarxiv-cs-ai
30 Jun 2026
Safety

APPO: Agentic Procedural Policy Optimization

DGX agent

arXiv:2606.12384v1 Announce Type: cross Abstract: Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabilities of large language model agents

safetyarxiv-cs-ai
11 Jun 2026
Model Releases

Causal Agent Replay: Counterfactual Attribution for LLM-Agent Failures

DGX agent

arXiv:2606.08275v1 Announce Type: cross Abstract: When an LLM agent fails -- issues a refund it should not have, calls the wrong tool, leaks data -- existing tooling answers what happened (observabili

model-releasesarxiv-cs-ai
9 Jun 2026
← Previous
1…1112131415…108
Next →