AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
5,141 results
Model Releases

MCP-Persona: Benchmarking LLM Agents on Real-World Personal Applications via Environment Simulation

DGX agent

arXiv:2606.02470v1 Announce Type: new Abstract: The Model Context Protocol (MCP) has emerged as a transformative standard for connecting large language models (LLMs) with external data sources and too

model-releasesarxiv-cs-ai
2 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say

DGX agent

arXiv:2606.00152v1 Announce Type: cross Abstract: LLM-based agents are rapidly advancing, autonomously invoking external tools to complete multi-step tasks for users. However, agents often acquire mor

model-releasesarxiv-cs-ai
2 Jun 2026
Safety

SPADER: Step-wise Peer Advantage with Diversity-Aware Exploration Rewards for Multi-Answer Question Answering

DGX agent

arXiv:2606.00593v1 Announce Type: cross Abstract: Large language models are increasingly deployed as tool-augmented agents to acquire information beyond parametric knowledge. While recent work has imp

safetyarxiv-cs-ai
2 Jun 2026
Model Releases

From Prompt Injection to Persistent Control: Defending Agentic Harness Against Trojan Backdoors

DGX agent

arXiv:2605.31042v1 Announce Type: cross Abstract: LLM agents are evolving from conversational chatbots to operational tools in real-world workspaces. In local agentic harnesses, an LLM can read and wr

model-releasesarxiv-cs-ai
1 Jun 2026
Agents

Governing Technical Debt in Agentic AI Systems

DGX agent

arXiv:2605.29129v1 Announce Type: new Abstract: Agentic AI systems are increasingly being explored as production infrastructure: they reason over multiple steps, call tools, act through workflows, and

agentsarxiv-cs-ai
29 May 2026
Model Releases

Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows

DGX agent

arXiv:2605.27922v1 Announce Type: new Abstract: LLM agents are increasingly deployed as executable systems that use tools, modify workspaces, and produce concrete artifacts. In such workflows, perform

model-releasesarxiv-cs-ai
28 May 2026
Local Ai

HARP: Measuring Harm Amplification in Multi-Agent LLM Systems

DGX agent

arXiv:2605.27489v1 Announce Type: cross Abstract: Multi-agent LLM systems decompose workflows across agents, tools, shared context, memory, and decision gates. This modularity improves interpretabilit

local-aiarxiv-cs-ai
28 May 2026
Agents

SPEAR: Code-Augmented Agentic Prompt Optimization

DGX agent

arXiv:2605.26275v1 Announce Type: new Abstract: Automatic prompt engineering (APE) rewrites prompts to improve downstream task performance, but existing APE loops treat the optimizer itself as a fixed

agentsarxiv-cs-cl
27 May 2026
Local Ai

The Constraint Tax: Measuring Validity-Correctness Tradeoffs in Structured Outputs for Small Language Models

DGX agent

arXiv:2605.26128v1 Announce Type: new Abstract: Production LLM systems increasingly require machine-readable outputs: JSON objects, typed traces, regex-constrained fields, and tool-call schemas. This

local-aiarxiv-cs-lg
27 May 2026
Applications

UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action

DGX agent

arXiv:2510.17790v3 Announce Type: replace-cross Abstract: Computer-use agents face a fundamental limitation. They rely exclusively on primitive GUI actions (click, type, scroll), creating brittle exec

applicationsarxiv-cs-cl
27 May 2026
Agents

ACC: Compiling Agent Trajectories for Long-Context Training

DGX agent

arXiv:2605.21850v1 Announce Type: new Abstract: Recent development of agents has renewed demand for long-context reasoning capacity of LLMs. However, training LLMs for this capacity requires costly lo

agentsarxiv-cs-cl
22 May 2026
Safety

Governance by Design: Architecting Agentic AI for Organizational Learning and Scalable Autonomy

DGX agent

arXiv:2605.20210v1 Announce Type: cross Abstract: Agentic AI systems - systems that can pursue goals through multi-step planning and tool-mediated action with limited direct supervision - are moving f

safetyarxiv-cs-ai
22 May 2026
Agents

OPERA: An Agent for Image Restoration with End-to-End Joint Planning-Execution Optimization

DGX agent

arXiv:2605.22104v1 Announce Type: new Abstract: Real-world image restoration is challenging due to complex and interacting mixed degradations. Recent agent-based approaches address this problem by com

agentsarxiv-cs-cv
22 May 2026
Model Releases

Teaching AI Through Benchmark Construction: QuestBench as a Course-Based Practice for Accountable Knowledge Work

DGX agent

arXiv:2605.21413v2 Announce Type: new Abstract: As AI becomes part of everyday learning, many courses teach students to use it mainly as a productivity tool: how to prompt, search, summarize, write, c

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

AgentAtlas: Beyond Outcome Leaderboards for LLM Agents

DGX agent

arXiv:2605.20530v1 Announce Type: cross Abstract: Large language model agents now act on codebases, browsers, operating systems, calendars, files, and tool ecosystems, but the benchmarks used to evalu

model-releasesarxiv-cs-cl
21 May 2026
Model Releases

RecoAtlas: From Semantic Plausibility to Set-Level Utility in LLM Recommendation Agents

DGX agent

arXiv:2605.18805v1 Announce Type: cross Abstract: LLM recommendation agents increasingly produce structured recommendation reports: sets of items accompanied by natural-language justifications. Yet ex

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

ContractBench: Can LLM Agents Preserve Observation Contracts?

DGX agent

arXiv:2605.17281v1 Announce Type: cross Abstract: Tool-augmented LLM agents call APIs whose intermediate outputs, such as presigned URLs, session tokens, and OAuth state parameters, are observation co

model-releasesarxiv-cs-ai
19 May 2026
Safety

Helpful to a Fault: Measuring Illicit Assistance in Multi-Turn, Multilingual LLM Agents

DGX agent

arXiv:2602.16346v3 Announce Type: replace Abstract: LLM-based agents execute real-world workflows via tools and memory. These affordances enable ill-intended adversaries to also use these agents to ca

safetyarxiv-cs-cl
19 May 2026
Model Releases

Homoglyph-based Adversarial Perturbation of Introductory Computer Science Theory Problems

DGX agent

arXiv:2605.16286v1 Announce Type: cross Abstract: Different AI tools such as ChatGPT, Gemini, and Claude are becoming very popular. Although they are helpful for many day-to-day tasks, they can be use

model-releasesarxiv-cs-ai
19 May 2026
Local Ai

Memisis: Orchestrating and Evaluating Synthetic Data for Tabular Health Datasets

DGX agent

arXiv:2605.17758v1 Announce Type: new Abstract: Synthetic data is widely used in healthcare to create datasets that are similar to original data but without the privacy concerns. Generating and evalua

local-aiarxiv-cs-lg
19 May 2026
Research

To Trust or Not to Trust: Authors' Response to AI-based Reviews

DGX agent

arXiv:2605.16623v1 Announce Type: cross Abstract: Large language models are increasingly discussed and used as tools that may assist with scholarly peer review, but empirical evidence regarding how au

researcharxiv-cs-ai
19 May 2026
Model Releases

CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing

DGX agent

arXiv:2605.14084v1 Announce Type: cross Abstract: Code agents must both reason over long-horizon repository state and obey strict tool-use protocols. In paired Instruct/Thinking checkpoints, these cap

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Good to Go: The LOOP Skill Engine That Hits 99% Success and Slashes Token Usage by 99% via One-Shot Recording and Deterministic Replay

DGX agent

arXiv:2605.14237v1 Announce Type: new Abstract: Deploying AI agents for repetitive periodic tasks exposes a critical tension: Large Language Models (LLMs) offer unmatched flexibility in tool orchestra

model-releasesarxiv-cs-ai
15 May 2026
Research

The Bicameral Model: Bidirectional Hidden-State Coupling Between Parallel Language Models

DGX agent

arXiv:2605.11167v1 Announce Type: new Abstract: Existing multi-model and tool-augmented systems communicate by generating text, serializing every exchange through the output vocabulary. Can two pretra

researcharxiv-cs-cl
13 May 2026
Agents

GenCellAgent: Generalizable, Training-Free Cellular Image Segmentation via Large Language Model Agents

DGX agent

arXiv:2510.13896v2 Announce Type: replace-cross Abstract: Cellular image segmentation is essential for quantitative biology yet remains difficult due to heterogeneous modalities, morphological variabi

agentsarxiv-cs-ai
12 May 2026
Model Releases

MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments

DGX agent

arXiv:2605.09131v1 Announce Type: new Abstract: The Model Context Protocol (MCP) has unified the interface between Large Language Models (LLMs) and external tools, yet a fundamental gap remains in how

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM Agents

DGX agent

arXiv:2605.08876v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous agents that execute tool-augmented, multi-step tasks, where latency is a critical f

model-releasesarxiv-cs-lg
12 May 2026
Agents

Willful Disobedience: Automatically Detecting Failures in Agentic Traces

DGX agent

arXiv:2603.23806v2 Announce Type: replace-cross Abstract: AI agents are increasingly embedded in real software systems, where they execute multi-step workflows through multi-turn dialogue, tool invoca

agentsarxiv-cs-ai
12 May 2026
Agents

Towards Security-Auditable LLM Agents: A Unified Graph Representation

DGX agent

arXiv:2605.06812v1 Announce Type: new Abstract: LLM-based agentic systems are rapidly evolving to perform complex autonomous tasks through dynamic tool invocation, stateful memory management, and mult

agentsarxiv-cs-ai
11 May 2026
Agents

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning

DGX agent

arXiv:2605.02913v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a central post-training tool for improving the reasoning abilities of large language models (LLMs). In these syst

agentsarxiv-cs-lg
6 May 2026
Model Releases

Parameter Space Analysis through Guided Visual Interpolations

DGX agent

arXiv:2509.19202v2 Announce Type: replace-cross Abstract: We propose Parameter Space Analysis through Guided Visual Interpolations (ParamInter), a novel tool for high-dimensional input parameter space

model-releasesarxiv-cs-lg
5 May 2026
Model Releases

The Compliance Gap: Why AI Systems Promise to Follow Process Instructions but Don't

DGX agent

arXiv:2605.01771v1 Announce Type: new Abstract: An auditor instructs an AI assistant: 'open each file individually using the Read tool -- no scripts, no agents.' The AI replies 'Yes' -- then issues a

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

Enforcing Benign Trajectories: A Behavioral Firewall for Structured-Workflow AI Agents

DGX agent

arXiv:2604.26274v1 Announce Type: cross Abstract: Structured-workflow agents driven by large language models execute tool calls against sensitive external environments. We propose odename, a telemetry

model-releasesarxiv-cs-ai
30 Apr 2026
Safety

Quantifying Divergence in Inter-LLM Communication Through API Retrieval and Ranking

DGX agent

arXiv:2604.22760v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly operate as autonomous agents that reason over external APIs to perform complex tasks. However, their reliabi

safetyarxiv-cs-ai
28 Apr 2026
Safety

SemML 2.0: Synthesizing Controllers for LTL

DGX agent

arXiv:2604.24102v1 Announce Type: new Abstract: Synthesizing a reactive system from specifications given in linear temporal logic (LTL) is a classical problem, finding its applications in safety-criti

safetyarxiv-cs-ai
28 Apr 2026
Tutorials

Pre-trained Large Language Models Learn Hidden Markov Models In-context

DGX agent

arXiv:2506.07298v3 Announce Type: replace-cross Abstract: Hidden Markov Models (HMMs) are foundational tools for modeling sequential data with latent Markovian structure, yet fitting them to real-worl

tutorialsarxiv-cs-ai
27 Apr 2026
Safety

Survey on Evaluation of LLM-based Agents

DGX agent

arXiv:2503.16416v2 Announce Type: replace Abstract: LLM-based agents represent a paradigm shift in AI, enabling autonomous systems to plan, reason, and use tools while interacting with dynamic environ

safetyarxiv-cs-ai
24 Apr 2026
Research

R2IF: Aligning Reasoning with Decisions via Composite Rewards for Interpretable LLM Function Calling

DGX agent

arXiv:2604.20316v1 Announce Type: new Abstract: Function calling empowers large language models (LLMs) to interface with external tools, yet existing RL-based approaches suffer from misalignment betwe

researcharxiv-cs-lg
23 Apr 2026
Agents

DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents

DGX agent

arXiv:2604.19264v1 Announce Type: new Abstract: Agentic multimodal models have garnered significant attention for their ability to leverage external tools to tackle complex tasks. However, it is obser

agentsarxiv-cs-cv
22 Apr 2026
Model Releases

PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts

DGX agent

arXiv:2506.06211v2 Announce Type: replace-cross Abstract: Puzzlehunts are a genre of complex, multi-step puzzles lacking well-defined problem definitions. In contrast to conventional reasoning benchma

model-releasesarxiv-cs-ai
22 Apr 2026
Agents

Whispers in the Machine: Confidentiality in Agentic Systems

DGX agent

arXiv:2402.06922v5 Announce Type: replace-cross Abstract: Large language model (LLM)-based agents combine LLMs with external tools to automate tasks such as scheduling meetings, managing documents, or

agentsarxiv-cs-lg
22 Apr 2026
Agents

GenericAgent: A Token-Efficient Self-Evolving LLM Agent via Contextual Information Density Maximization (V1.0)

DGX agent

arXiv:2604.17091v1 Announce Type: new Abstract: Long-horizon large language model (LLM) agents are fundamentally limited by context. As interactions become longer, tool descriptions, retrieved memorie

agentsarxiv-cs-cl
21 Apr 2026
Research

LLMbench: A Comparative Close Reading Workbench for Large Language Models

DGX agent

arXiv:2604.15508v1 Announce Type: cross Abstract: LLMbench is a browser-based workbench for the comparative close reading of large language model (LLM) outputs. Where existing tools for LLM comparison

researcharxiv-cs-ai
20 Apr 2026
Model Releases

SafeHarness: Lifecycle-Integrated Security Architecture for LLM-based Agent Deployment

DGX agent

arXiv:2604.13630v1 Announce Type: cross Abstract: The performance of large language model (LLM) agents depends critically on the execution harness, the system layer that orchestrates tool use, context

model-releasesarxiv-cs-ai
17 Apr 2026
Research

The Code Whisperer: LLM and Graph-Based AI for Smell and Vulnerability Resolution

DGX agent

arXiv:2604.13114v1 Announce Type: cross Abstract: Code smells and software vulnerabilities both increase maintenance cost, yet they are often handled by separate tools that miss structural context and

researcharxiv-cs-ai
17 Apr 2026
Safety

Beyond Static Sandboxing: Learned Capability Governance for Autonomous AI Agents

DGX agent

arXiv:2604.11839v1 Announce Type: cross Abstract: Autonomous AI agents built on open-source runtimes such as OpenClaw expose every available tool to every session by default, regardless of the task. A

safetyarxiv-cs-ai
15 Apr 2026
Agents

El Agente Quntur: A research collaborator agent for quantum chemistry

DGX agent

arXiv:2602.04850v2 Announce Type: replace-cross Abstract: Quantum chemistry is a foundational enabling tool for the fields of chemistry, materials science, computational biology and others. Despite of

agentsarxiv-cs-ai
15 Apr 2026
Applications

LLM-Based Automated Diagnosis Of Integration Test Failures At Google

DGX agent

arXiv:2604.12108v1 Announce Type: cross Abstract: Integration testing is critical for the quality and reliability of complex software systems. However, diagnosing their failures presents significant c

applicationsarxiv-cs-ai
15 Apr 2026
← Previous
1…1415161718…108
Next →