AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
5,141 results
Agents

MemToolAgent overview with a simple restaurant booking scenario where the agent retrieves similar memories, receives feedback on an invalid time format, and generates a reflection to update its memory

DGX agent

arXiv:2606.07909v1 Announce Type: new Abstract: Modern large language model (LLM) agents can use external tools to help users solve complex tasks. However, for problems that require learning from long

agentsarxiv-cs-ai
9 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Agents

Projecting the Emerging Mindset of SWE Agent by Launching a Wild Code Understanding Journey

DGX agent

arXiv:2606.08500v1 Announce Type: cross Abstract: Software engineering agents (SWE agents) increasingly work through tool-mediated trajectories in real repositories, yet their behavior remains difficu

agentsarxiv-cs-ai
9 Jun 2026
Safety

SciTrace: Trajectory-Aware Safety Reasoning for Scientific Discovery Agents

DGX agent

arXiv:2606.08234v1 Announce Type: new Abstract: LLM-based scientific agents have shown strong capacity for autonomous research, yet their safety layers remain structurally divorced from core reasoning

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents

DGX agent

arXiv:2606.04874v1 Announce Type: new Abstract: Planning is central to LLM agents: before acting, an agent must decompose goals, select tools, reason over constraints, and decide when a task is infeas

model-releasesarxiv-cs-cl
4 Jun 2026
Agents

Description-Code Inconsistency in Real-world MCP Servers: Measurement, Detection, and Security Implications

DGX agent

arXiv:2606.04769v1 Announce Type: cross Abstract: The Model Context Protocol (MCP) has emerged as a critical standard empowering Large Language Models (LLMs) to utilize external tools. In this ecosyst

agentsarxiv-cs-ai
4 Jun 2026
Model Releases

AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations

DGX agent

arXiv:2606.02240v1 Announce Type: cross Abstract: Indirect prompt injection in tool-use agents is a concrete production threat: LLM agents read from integrations (third-party services such as Gmail, S

model-releasesarxiv-cs-ai
2 Jun 2026
Safety

AIRGuard: Guarding Agent Actions with Runtime Authority Control

DGX agent

arXiv:2605.28914v1 Announce Type: cross Abstract: Tool-using language agents turn model decisions into external side effects: they read files, run scripts, call APIs, send messages, and invoke Model C

safetyarxiv-cs-ai
29 May 2026
Applications

It`s All About Speed: AI`s Impact on Workflow in Music Production

DGX agent

arXiv:2605.29931v1 Announce Type: new Abstract: In this paper, we present the results of an ethnographic study into the impact of AI and automated tools on music production workflow. Focusing specific

applicationsarxiv-cs-ai
29 May 2026
Model Releases

Scaling Laws for Agent Harnesses via Effective Feedback Compute

DGX agent

arXiv:2605.29682v1 Announce Type: new Abstract: Agent harnesses increasingly determine the performance of language-model systems by deciding how models call tools, receive feedback, verify intermediat

model-releasesarxiv-cs-cl
29 May 2026
Tutorials

An investigation of AI integration in sound designer workflows and experiences

DGX agent

arXiv:2605.27174v1 Announce Type: cross Abstract: Artificial intelligence is increasingly being integrated into professional audio production workflows, yet a gap persists between the tools developers

tutorialsarxiv-cs-ai
27 May 2026
Model Releases

ConVer: Using Contracts and Loop Invariant Synthesis for Scalable Formal Software Verification

DGX agent

arXiv:2605.27051v1 Announce Type: cross Abstract: Formal verification of large C programs is impeded by state-space explosion: Bounded Model Checking (BMC) tools must encode the entire state space up

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning

DGX agent

arXiv:2605.25200v1 Announce Type: new Abstract: Travel planning is a realistic task for evaluating the planning and tool-use abilities of LLM agents. However, existing benchmarks typically assume only

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

When the Manual Lies: A Realistic Benchmark to Evaluate MCP Poisoning Attacks for LLM Agents

DGX agent

arXiv:2605.24069v1 Announce Type: cross Abstract: The rise of tool-using Large Language Model (LLM) agents, standardized by protocols like the Model Context Protocol (MCP), has unlocked unprecedented

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI

DGX agent

arXiv:2603.14987v2 Announce Type: replace Abstract: Agentic AI systems increasingly act through tool-augmented, multi-step workflows whose failures (unsafe tool use, unauthorised actions, social harm)

model-releasesarxiv-cs-cl
22 May 2026
Safety

Learning to Configure Agentic AI Systems

DGX agent

arXiv:2602.11574v3 Announce Type: replace Abstract: Configuring LLM-based agent systems involves choosing workflows, tools, token budgets, and prompts from a large combinatorial design space, and is t

safetyarxiv-cs-ai
22 May 2026
Agents

Agent JIT Compilation for Latency-Optimizing Web Agent Planning and Scheduling

DGX agent

arXiv:2605.21470v1 Announce Type: new Abstract: Computer-use agents (CUA) automate tasks specified with natural language such as 'order the cheapest item from Taco Bell' by generating sequences of cal

agentsarxiv-cs-lg
21 May 2026
Safety

Progent: Securing AI Agents with Privilege Control

DGX agent

arXiv:2504.11703v3 Announce Type: replace-cross Abstract: AI agents interact with external environments through tool calls, exposing them to attacks like indirect prompt injection that can trigger una

safetyarxiv-cs-ai
15 May 2026
Model Releases

SupChain-Bench: Benchmarking Large Language Models for Real-World Supply Chain Management

DGX agent

arXiv:2602.07342v2 Announce Type: replace Abstract: Large language models (LLMs) have shown promise in complex reasoning and tool-based decision making, motivating their application to real-world supp

model-releasesarxiv-cs-ai
14 May 2026
Model Releases

Oracle Poisoning: Corrupting Knowledge Graphs to Weaponise AI Agent Reasoning

DGX agent

arXiv:2605.09822v1 Announce Type: cross Abstract: We define Oracle Poisoning, an attack class in which an adversary corrupts a structured knowledge graph that AI agents query at runtime via tool-use p

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents

DGX agent

arXiv:2605.07177v1 Announce Type: cross Abstract: Existing multimodal search agents process target entities sequentially, issuing one tool call per entity and accumulating redundant interaction rounds

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live

DGX agent

arXiv:2511.02230v4 Announce Type: replace-cross Abstract: KV cache management is essential for efficient LLM inference. To maximize utilization, existing inference engines evict finished requests' KV

model-releasesarxiv-cs-ai
7 May 2026
Applications

Human-AI Co-Mentorship in Project-Based Learning: A Case Study in Financial Forecasting

DGX agent

arXiv:2605.05144v1 Announce Type: new Abstract: This paper reflects on a AI research project carried out by a team of high-school and early-undergraduate students under the mentorship of graduate rese

applicationsarxiv-cs-lg
7 May 2026
Agents

Separating Intelligence from Execution: A Workflow Engine for the Model Context Protocol

DGX agent

arXiv:2605.00827v1 Announce Type: cross Abstract: Large Language Model (LLM) agents increasingly interact with external systems through tool-calling protocols such as the Model Context Protocol (MCP).

agentsarxiv-cs-ai
6 May 2026
Model Releases

Towards Multi-Agent Autonomous Reasoning in Hydrodynamics

DGX agent

arXiv:2605.01102v1 Announce Type: new Abstract: Single-agent systems (SAS) have become the default pattern for LLM-driven scientific workflows, but routing planning, tool use, and synthesis through a

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Agentic Education: Using Claude Code to Teach Claude Code

DGX agent

arXiv:2604.17460v2 Announce Type: replace-cross Abstract: AI coding assistants have proliferated rapidly, yet structured pedagogical frameworks for learning these tools remain scarce. Developers face

model-releasesarxiv-cs-ai
1 May 2026
Applications

BAss: Symbolic Reasoning in Abstract Dialectical Frameworks

DGX agent

arXiv:2604.27576v1 Announce Type: cross Abstract: We present BAss (BDD-based ADF symbolic solver), a novel analysis tool for Abstract Dialectical Frameworks (ADFs) based on Binary Decision Diagrams (B

applicationsarxiv-cs-lg
1 May 2026
Applications

IACDM: Interactive Adversarial Convergence Development Methodology -- A Structured Framework for AI-Assisted Software Development

DGX agent

arXiv:2604.16399v2 Announce Type: replace-cross Abstract: The widespread adoption of AI-assisted development tools in 2025 -- and the emergence of vibe coding, a practice of generating complete applic

applicationsarxiv-cs-ai
1 May 2026
Safety

The Likelihood Ratio Wall: Structural Limits on Accurate Risk Assessment for Rare Violence

DGX agent

arXiv:2604.27282v1 Announce Type: cross Abstract: Pretrial risk assessment tools are used on over one million U.S. defendants each year, yet their use for predicting rare violent re-offense faces a ba

safetyarxiv-cs-lg
1 May 2026
Tutorials

Upskilling with Generative AI: Practices and Challenges for Freelance Knowledge Workers

DGX agent

arXiv:2604.27231v1 Announce Type: cross Abstract: Freelance workers must continually acquire new skills to remain competitive in online labor markets, yet they lack the organizational training, mentor

tutorialsarxiv-cs-ai
1 May 2026
Agents

Biomedical systems biology workflow orchestration and execution with PoSyMed

DGX agent

arXiv:2604.20906v1 Announce Type: cross Abstract: The rapid growth of scientific software has created practical barriers for bioinformatics research. Although powerful statistical, artificial intellig

agentsarxiv-cs-ai
24 Apr 2026
Safety

Rethinking Reinforcement Fine-Tuning in LVLM: Convergence, Reward Decomposition, and Generalization

DGX agent

arXiv:2604.19857v1 Announce Type: cross Abstract: Reinforcement fine-tuning with verifiable rewards (RLVR) has emerged as a powerful paradigm for equipping large vision-language models (LVLMs) with ag

safetyarxiv-cs-cl
23 Apr 2026
Model Releases

When Safety Fails Before the Answer: Benchmarking Harmful Behavior Detection in Reasoning Chains

DGX agent

arXiv:2604.19001v1 Announce Type: new Abstract: Large reasoning models (LRMs) produce complex, multi-step reasoning traces, yet safety evaluation remains focused on final outputs, overlooking how harm

model-releasesarxiv-cs-cl
22 Apr 2026
Tutorials

A Comparative Study on the Impact of Traditional Learning and Interactive Learning on Students' Academic Performance and Emotional Well-Being

DGX agent

arXiv:2604.15335v1 Announce Type: cross Abstract: The growing adoption of interactive learning tools in higher education offers new opportunities to enhance student performance and well-being. This st

tutorialsarxiv-cs-ai
20 Apr 2026
Agents

Capture the Flags: Family-Based Evaluation of Agentic LLMs via Semantics-Preserving Transformations

DGX agent

arXiv:2602.05523v2 Announce Type: replace-cross Abstract: Agentic large language models (LLMs) are increasingly evaluated on cybersecurity tasks using capture-the-flag (CTF) benchmarks, yet existing p

agentsarxiv-cs-ai
20 Apr 2026
Applications

Joint Score-Threshold Optimization for Interpretable Risk Assessment

DGX agent

arXiv:2510.21934v3 Announce Type: replace Abstract: Risk assessment tools in healthcare commonly employ point-based scoring systems that map patients to ordinal risk categories via thresholds. While e

applicationsarxiv-cs-lg
20 Apr 2026
Safety

IROSA: Interactive Robot Skill Adaptation using Natural Language

DGX agent

arXiv:2603.03897v3 Announce Type: replace-cross Abstract: Foundation models have demonstrated impressive capabilities across diverse domains, while imitation learning provides principled methods for r

safetyarxiv-cs-cl
17 Apr 2026
Safety

Gradient boundaries through confidence intervals for forced alignment estimates using model ensembles

DGX agent

arXiv:2506.01256v4 Announce Type: replace-cross Abstract: Forced alignment is a common tool to align audio with orthographic and phonetic transcriptions. Most forced alignment tools provide only point

safetyarxiv-cs-cl
15 Apr 2026
Applications

Perceived Importance of Cognitive Skills Among Computing Students in the Era of AI

DGX agent

arXiv:2604.10730v1 Announce Type: cross Abstract: The availability and increasing integration of generative AI tools have transformed computing education. While AI in education presents opportunities,

applicationsarxiv-cs-ai
14 Apr 2026
Research

Towards an Appropriate Level of Reliance on AI: A Preliminary Reliance-Control Framework for AI in Software Engineering

DGX agent

arXiv:2604.10530v1 Announce Type: cross Abstract: How software developers interact with Artificial Intelligence (AI)-powered tools, including Large Language Models (LLMs), plays a vital role in how th

researcharxiv-cs-ai
14 Apr 2026
Model Releases

ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis

DGX agent

arXiv:2604.02022v2 Announce Type: replace Abstract: Evaluating the safety of LLM-based agents is increasingly important because risks in realistic deployments often emerge over multi-step interactions

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

Mitigating Distribution Sharpening in Math RLVR via Distribution-Aligned Hint Synthesis and Backward Hint Annealing

DGX agent

arXiv:2604.07747v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) can improve low-k reasoning accuracy while narrowing solution coverage on challenging math que

model-releasesarxiv-cs-cl
10 Apr 2026
Model Releases

DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?

DGX agent

arXiv:2608.10366v1 Announce Type: new Abstract: Real-world data science involves long-horizon workflows that span data wrangling, exploration, modeling, visualization, and validation, and require coor

model-releasesarxiv-cs-ai
12 Aug 2026
Agents

On Understanding, Identifying, and Mitigating Vulnerabilities in Agentic Large Language Models

DGX agent

arXiv:2608.10530v1 Announce Type: cross Abstract: Large Language Models (LLMs) have undergone a shift from stateless conversational interfaces to autonomous agents capable of multi-step planning, tool

agentsarxiv-cs-ai
12 Aug 2026
Model Releases

REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems

DGX agent

arXiv:2608.10669v1 Announce Type: new Abstract: Large language model (LLM) agents combine language-based reasoning with external tools to perform complex tasks. Adversarial inputs can exploit interact

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

Toward Metacognitive One-Shot Indirect Prompt Injection: Strategy Abstraction Via Outcome-Conditioned Reflection

DGX agent

arXiv:2608.08795v1 Announce Type: cross Abstract: Tool-using large language model (LLM) agents are vulnerable to indirect prompt injection (IPI), in which malicious instructions embedded in external o

model-releasesarxiv-cs-cl
11 Aug 2026
Agents

VDGR-RAG: Vectors, Directories, Graphs, and Reflection Are All You Need for Unified Reasoning over Hierarchical Enterprise Knowledge

DGX agent

arXiv:2608.07994v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) is essential for enterprise knowledge question answering (QA), particularly in domains with complex product documen

agentsarxiv-cs-ai
11 Aug 2026
Model Releases

ZhuLong: Execution-Grounded LLM Agent for EDA Scripting with Offline API Self-Exploration

DGX agent

arXiv:2608.07925v1 Announce Type: new Abstract: EDA scripting with tool-specific, often undocumented APIs remains a long-tail bottleneck that existing LLMs fail to address. This paper presents ZhuLong

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Who Checks the Citations? Benchmarking Legal Hallucination Detection

DGX agent

arXiv:2606.21155v2 Announce Type: replace Abstract: Attorneys, judges, and pro se filers increasingly use AI to draft legal documents, yet these tools frequently fabricate citations. Despite predictio

model-releasesarxiv-cs-cl
7 Aug 2026
← Previous
1…1213141516…108
Next →