AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
5,141 results
Research

ActiNet: An Open-Source Tool for Activity Intensity Classification of Wrist-Worn Accelerometry Using Self-Supervised Deep Learning

DGX agent

arXiv:2510.01712v2 Announce Type: replace Abstract: The use of accurate and reliable open-source human activity recognition (HAR) models on passively collected wrist-accelerometer data is essential in

researcharxiv-cs-lg
1 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Applications

Persuadability and LLMs as Legal Decision Tools

DGX agent

arXiv:2604.26233v1 Announce Type: new Abstract: As Large Language Models (LLMs) are proposed as legal decision assistants, and even first-instance decision-makers, across a range of judicial and admin

applicationsarxiv-cs-ai
30 Apr 2026
Tutorials

BridgeACT: Bridging Human Demonstrations to Robot Actions via Unified Tool-Target Affordances

DGX agent

arXiv:2604.23249v1 Announce Type: new Abstract: Learning robot manipulation from human videos is appealing due to the scale and diversity of human demonstrations, but transferring such demonstrations

tutorialsarxiv-cs-ro
28 Apr 2026
Safety

Why Do Language Model Agents Whistleblow?

DGX agent

arXiv:2511.17085v3 Announce Type: replace-cross Abstract: The deployment of Large Language Models (LLMs) as tool-using agents causes their alignment training to manifest in new ways. Recent work finds

safetyarxiv-cs-ai
24 Apr 2026
Agents

LLM as a Tool, Not an Agent: Code-Mined Tree Transformations for Neural Architecture Search

DGX agent

arXiv:2604.16555v1 Announce Type: cross Abstract: Neural Architecture Search (NAS) aims to automatically discover high-performing deep neural network (DNN) architectures. However, conventional algorit

agentsarxiv-cs-cv
21 Apr 2026
Research

Toward Efficient Influence Function: Dropout as a Compression Tool

DGX agent

arXiv:2509.15651v2 Announce Type: replace Abstract: Assessing the impact the training data on machine learning models is crucial for understanding the behavior of the model, enhancing the transparency

researcharxiv-cs-lg
21 Apr 2026
Model Releases

Applying an Agentic Coding Tool for Improving Published Algorithm Implementations

DGX agent

arXiv:2604.13109v1 Announce Type: cross Abstract: We present a two-stage pipeline for AI-assisted improvement of published algorithm implementations. In the first stage, a large language model with re

model-releasesarxiv-cs-ai
17 Apr 2026
Research

MyoVision: A Mobile Research Tool and NEATBoost-Attention Ensemble Framework for Real Time Chicken Breast Myopathy Detection

DGX agent

arXiv:2604.13456v1 Announce Type: cross Abstract: Woody Breast (WB) and Spaghetti Meat (SM) myopathies significantly impact poultry meat quality, yet current detection methods rely either on subjectiv

researcharxiv-cs-cv
16 Apr 2026
Model Releases

FlexMS is a flexible framework for benchmarking deep learning-based mass spectrum prediction tools in metabolomics

DGX agent

arXiv:2602.22822v2 Announce Type: replace Abstract: The identification and property prediction of chemical molecules is of central importance in the advancement of drug discovery and material science,

model-releasesarxiv-cs-ai
14 Apr 2026
Safety

ToolUniverse: An open platform for democratizing AI scientists

DGX agent

arXiv:2509.23426v3 Announce Type: replace Abstract: AI scientists are emerging computational systems that serve as collaborative partners in discovery. These systems remain difficult to build because

safetyarxiv-cs-ai
11 Aug 2026
Agents

OTAP: Structure-Aware Optimal Transport for Evaluating Planning and Execution in Agent Trajectories

DGX agent

arXiv:2607.17082v2 Announce Type: replace-cross Abstract: Large language model agents solve tasks by generating trajectories that interleave planning, tool calls, and intermediate results. Current eva

agentsarxiv-cs-cl
4 Aug 2026
Model Releases

ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents

DGX agent

arXiv:2606.18037v2 Announce Type: replace Abstract: Tool-using LLM agents increasingly use the Model Context Protocol (MCP) to answer from heterogeneous evidence sources, including search, APIs, datab

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

DeepTravel: An End-to-End Agentic Reinforcement Learning Framework for Autonomous Travel Planning Agents

DGX agent

arXiv:2509.21842v2 Announce Type: replace Abstract: Travel planning (TP) agent has recently worked as an emerging building block to interact with external tools/resources for travel itinerary generati

model-releasesarxiv-cs-ai
15 Jul 2026
Safety

TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration

DGX agent

arXiv:2603.27742v2 Announce Type: replace Abstract: Vision-language agents that orchestrate specialized tools for image restoration (IR) have emerged as a promising method, yet most existing framework

safetyarxiv-cs-cv
9 Jul 2026
Model Releases

XSkill: Continual Learning from Experience and Skills in Multimodal Agents

DGX agent

arXiv:2603.12056v3 Announce Type: replace Abstract: Multimodal agents can now tackle complex reasoning tasks with diverse tools, yet they still suffer from inefficient tool use and inflexible orchestr

model-releasesarxiv-cs-ai
2 Jul 2026
Research

The Interplay of Harness Design and Post-Training in LLM Agents

DGX agent

arXiv:2606.25447v1 Announce Type: cross Abstract: Tool-integrated LLM agents are often wrapped within a harness: the scaffolding that determines which tools are exposed, how they are described, and wh

researcharxiv-cs-cl
25 Jun 2026
Model Releases

Verifiable Manifest Signing and Transparency Enforcement for Secure MCP-Based LLM Pipelines

DGX agent

arXiv:2601.23132v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly deployed in tool-driven environments such as healthcare analytics, financial systems, retrieval-

model-releasesarxiv-cs-ai
25 Jun 2026
Model Releases

The Surface You Test Is Not the Surface That Breaks

DGX agent

arXiv:2605.30454v1 Announce Type: cross Abstract: Tool-augmented LLM agents are vulnerable to prompt injection: a third party who controls part of the agent's context can plant instructions that the a

model-releasesarxiv-cs-ai
1 Jun 2026
Safety

BEAMS: Benchmarking and Evaluating AI for Modeling and Simulation

DGX agent

arXiv:2605.28994v1 Announce Type: new Abstract: AI tools to support real world decision making must be able to build simulation models that inform their recommendations and render them interpretable.

safetyarxiv-cs-ai
29 May 2026
Safety

Agent Explorative Policy Optimization for Multimodal Agentic Reasoning

DGX agent

arXiv:2605.28774v1 Announce Type: new Abstract: Vision-language models with extended reasoning succeed on complex problems, but many real-world problems require external tools that internal reasoning

safetyarxiv-cs-cl
28 May 2026
Model Releases

AI-Generated Slides: Are They Good? Can Students Tell?

DGX agent

arXiv:2605.13532v1 Announce Type: new Abstract: As generative AI (GenAI) tools become easily accessible, there is promise in using such tools to support instructors. To that end, this paper examines u

model-releasesarxiv-cs-ai
14 May 2026
Agents

Hybrid Inspection and Task-Based Access Control in Zero-Trust Agentic AI

DGX agent

arXiv:2605.02682v1 Announce Type: new Abstract: Authorizing Large Language Model (LLM)-driven agents to dynamically invoke tools and access protected resources introduces significant security risks, a

agentsarxiv-cs-ai
6 May 2026
Model Releases

Learning to Ask: When LLM Agents Meet Unclear Instruction

DGX agent

arXiv:2409.00557v4 Announce Type: replace-cross Abstract: Equipped with the capability to call functions, modern large language models (LLMs) can leverage external tools for addressing a range of task

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks

DGX agent

arXiv:2604.20087v1 Announce Type: new Abstract: Skills have become the de facto way to enable LLM agents to perform complex real-world tasks with customized instructions, workflows, and tools, but how

model-releasesarxiv-cs-cl
23 Apr 2026
Model Releases

How Adversarial Environments Mislead Agentic AI?

DGX agent

arXiv:2604.18874v1 Announce Type: new Abstract: Tool-integrated agents are deployed on the premise that external tools ground their outputs in reality. Yet this very reliance creates a critical attack

model-releasesarxiv-cs-ai
22 Apr 2026
Hardware

Beyond Visibility: Real-Time Surface Accessibility Fields from Sparse LiDAR

DGX agent

arXiv:2608.06412v1 Announce Type: cross Abstract: Understanding which surfaces in a scene are physically accessible to a given tool is fundamental for robotic interaction, yet 3D perception systems ty

hardwarearxiv-cs-ro
10 Aug 2026
Local Ai

CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video Understanding

DGX agent

arXiv:2607.24582v1 Announce Type: cross Abstract: Long-video understanding increasingly relies on large vision-language models and tool-augmented reasoning, but most systems apply the same inference p

local-aiarxiv-cs-ai
28 Jul 2026
Safety

Execution-Grounded Security Testing for Coding Agents in Software Engineering Pipelines

DGX agent

arXiv:2607.22569v1 Announce Type: new Abstract: Coding agents are increasingly integrated into system operations, where their tool use can directly modify project artifacts, execution environments, an

safetyarxiv-cs-ai
28 Jul 2026
Model Releases

Who Analyses the Analyser? Self-Validating LLM Hazard Analysis with Constitutional Meta-STPA

DGX agent

arXiv:2607.08054v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly trusted to draft the artifacts of safety analysis such as, losses, hazards, Unsafe Control Actions (UCAs

model-releasesarxiv-cs-ai
10 Jul 2026
Agents

Capability Gates Are Not Authorization: Confused-Deputy Failures in LLM Agent Frameworks

DGX agent

arXiv:2606.28679v1 Announce Type: cross Abstract: Tool-using LLM agents increasingly read untrusted content while holding side-effecting tools such as payments, email, CRM, and infrastructure APIs, ye

agentsarxiv-cs-ai
30 Jun 2026
Safety

RUBAS: Rubric-Based Reinforcement Learning for Agent Safety

DGX agent

arXiv:2606.04051v1 Announce Type: cross Abstract: The evolution of LLMs into tool-enabled agents creates a new class of safety challenges associated with real-world execution rather than simple text g

safetyarxiv-cs-ai
4 Jun 2026
Agents

Early Diagnosis of Wasted Computation in Multi-Agent LLM Systems via Failure-Aware Observability

DGX agent

arXiv:2606.01365v1 Announce Type: new Abstract: Tool-using multi-agent large language model (LLM) systems spend computation through model tokens, tool calls, retries, and code execution before produci

agentsarxiv-cs-ai
2 Jun 2026
Model Releases

Temporal Stability and Few-Shot Prompting in Math Task Assessment

DGX agent

arXiv:2605.30151v1 Announce Type: new Abstract: As AI tools become increasingly integrated into educational contexts, questions arise about both their stability over time and their responsiveness to p

model-releasesarxiv-cs-ai
29 May 2026
Safety

Governance by Construction for Generalist Agents

DGX agent

arXiv:2605.20874v1 Announce Type: new Abstract: Enterprise agents are increasingly expected to operate autonomously across tools and interfaces, yet production deployments require governance by constr

safetyarxiv-cs-ai
22 May 2026
Agents

SpecHop: Continuous Speculation for Accelerating Multi-Hop Retrieval Agents

DGX agent

arXiv:2605.21965v1 Announce Type: new Abstract: Large language models increasingly use external tools such as web search and document retrieval to solve information-intensive tasks. However, multi-hop

agentsarxiv-cs-cl
22 May 2026
Model Releases

Is Grep All You Need? How Agent Harnesses Reshape Agentic Search

DGX agent

arXiv:2605.15184v1 Announce Type: new Abstract: Recent advances in Large Language Model (LLM) agents have enabled complex agentic workflows where models autonomously retrieve information, call tools,

model-releasesarxiv-cs-cl
15 May 2026
Agents

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence

DGX agent

arXiv:2604.18292v1 Announce Type: cross Abstract: Large language models are increasingly expected to serve as general-purpose agents that interact with external, stateful tool environments. The Model

agentsarxiv-cs-cl
21 Apr 2026
Model Releases

GeoForge: Non-Parametric Self-Evolving Agents for Earth-Observation Reasoning

DGX agent

arXiv:2608.10494v1 Announce Type: new Abstract: Earth observation (EO) agents construct scientifically valid tool workflows and ground their conclusions in current geospatial evidence. This is challen

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

LLM within MCP Matters: Measuring Inefficient Resource Utilization Driven by LLMs

DGX agent

arXiv:2608.08467v1 Announce Type: new Abstract: The Model Context Protocol (MCP) standardizes how servers expose data and tools to Large Language Models (LLMs). A common server design embeds frequentl

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Winning by Peeking: Unenforced Budgets and Test-Set Selection Inflate Short-Budget AutoML Comparisons

DGX agent

arXiv:2608.07303v1 Announce Type: new Abstract: Comparisons between AutoML systems at short time budgets -- tens of seconds rather than hours -- are common in tool READMEs and workshop papers, and the

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning

DGX agent

arXiv:2607.26977v1 Announce Type: new Abstract: Travel planning is a demanding stress test for tool-using LLM agents: a usable itinerary is a single artifact that must be right along many axes at once

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic Evaluation

DGX agent

arXiv:2607.24772v1 Announce Type: new Abstract: Geoscience research requires complex analysis and domain expertise, with remote sensing (RS) observations as a key foundation. However, existing RS agen

model-releasesarxiv-cs-ai
29 Jul 2026
Local Ai

Auditing Provenance Sensitivity in LLM Agent Action Selection

DGX agent

arXiv:2607.20827v1 Announce Type: new Abstract: LLM agents choose tools and arguments from context that mixes user requests, tool outputs, retrieved records, memory, and untrusted text. Evidence can b

local-aiarxiv-cs-ai
24 Jul 2026
Model Releases

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making

DGX agent

arXiv:2607.20491v1 Announce Type: new Abstract: Standard evaluation benchmarks measure what a tool-using agent decides, not whether it arrives at that decision through the same process each time. We i

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

AgentCgroup: Understanding and Controlling OS Resources of AI Agents

DGX agent

arXiv:2602.09345v3 Announce Type: replace-cross Abstract: AI agents are increasingly deployed in multi-tenant cloud environments, where they execute diverse tool calls within sandboxed containers, eac

model-releasesarxiv-cs-ai
23 Jul 2026
Safety

Membership Inference Attacks for Unseen Classes

DGX agent

arXiv:2506.06488v3 Announce Type: replace Abstract: A key tool in developing safe AI models is data auditing, i.e., using statistical tools to determine whether harmful content may have been used in t

safetyarxiv-cs-lg
23 Jul 2026
Model Releases

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents

DGX agent

arXiv:2607.13988v1 Announce Type: new Abstract: Multi-turn agents solve complex tasks through extended sequences of tool interactions before producing a final answer, making credit assignment a fundam

model-releasesarxiv-cs-lg
16 Jul 2026
Model Releases

OmniGAIA: Towards Native Omni-Modal AI Agents

DGX agent

arXiv:2602.22897v3 Announce Type: replace Abstract: Human intelligence naturally intertwines omni-modal perception -- spanning vision, audio, and language -- with complex reasoning and tool usage to i

model-releasesarxiv-cs-ai
3 Jul 2026
← Previous
1…1011121314…108
Next →