AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,289 results
Safety

AttriMem: Attribution-Guided Process Feedback for Agent Memory Learning

DGX agent

arXiv:2607.21106v1 Announce Type: new Abstract: Effective memory is crucial for LLM agents, yet constructing it effectively remains challenging. A memory-construction policy decides what information t

safetyarxiv-cs-ai
24 Jul 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

CAMeR: Keyword-Gated Hybrid Activation for Adaptive Memory Retention in LLM Agents

DGX agent

arXiv:2607.20458v1 Announce Type: cross Abstract: Large language model (LLM) agents operating over extended dialogues accumulate vast amounts of information, yet existing memory systems either retain

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

CANN Bench: Benchmarking Agent Generated Kernels against Real NPU and Algorithmic Limits

DGX agent

arXiv:2607.20518v1 Announce Type: new Abstract: AI agents are now capable of writing, compiling, and iteratively optimizing low-level operator kernels on different hardware platforms. Existing benchma

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making

DGX agent

arXiv:2607.20491v1 Announce Type: new Abstract: Standard evaluation benchmarks measure what a tool-using agent decides, not whether it arrives at that decision through the same process each time. We i

model-releasesarxiv-cs-ai
24 Jul 2026
Safety

From Agent Failures to Text Policies: What Works and What Breaks

DGX agent

arXiv:2607.20668v1 Announce Type: cross Abstract: TextGrad improves language-model systems by revising text from feedback. Its core thesis is that natural-language feedback can act as a gradient for o

safetyarxiv-cs-ai
24 Jul 2026
Model Releases

GuardianAgentBench: Where Agents Fail and How to Guard Them

DGX agent

arXiv:2607.20982v1 Announce Type: new Abstract: As large language model agents increasingly operate autonomously with access to tools and external environments, ensuring their safe and reliable behavi

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

MemTools: A Unified Research Framework for Interoperable Agent Memory

DGX agent

arXiv:2607.21404v1 Announce Type: new Abstract: While memory systems are essential for agent architectures, pervasive architectural fragmentation restricts systematic research. Existing implementation

model-releasesarxiv-cs-cl
24 Jul 2026
Agents

MKEvolve: A Modular Multi-Agent Framework for Kernel Code Generation

DGX agent

arXiv:2607.20501v1 Announce Type: new Abstract: Despite rapid progress in LLM-based code generation, writing correct and performant kernels for hardware accelerators remains a key bottleneck in scalin

agentsarxiv-cs-ai
24 Jul 2026
Safety

Regulating autonomous and agentic AI

DGX agent

arXiv:2607.21345v1 Announce Type: new Abstract: Regulating activities where regulatees use autonomous and agentic AI is challenging. Regulatory assumptions about regulatee knowledge and control no lon

safetyarxiv-cs-ai
24 Jul 2026
Model Releases

Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs

DGX agent

arXiv:2509.14257v3 Announce Type: replace-cross Abstract: Large Language Model agents achieve strong performance on multi-step reasoning and tool-use tasks, but their impressive capabilities typically

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

ETPDesigner: Multi-Agent Orchestration for Interactive Multimodal Electronic Theater Program

DGX agent

arXiv:2607.19947v1 Announce Type: new Abstract: Electronic Theater Programs (ETPs) serve as critical promotional media in the performing arts, comprising a multi-page collection of heterogeneous visua

model-releasesarxiv-cs-cv
23 Jul 2026
Model Releases

NEXUS: Structured Runtime Safety for Tool-Using LLM Agents

DGX agent

arXiv:2607.19356v1 Announce Type: new Abstract: Tool-using LLM agents increasingly execute high-impact actions, making runtime safety monitoring essential. We present NEXUS (Neural EXecution Utility a

model-releasesarxiv-cs-ai
23 Jul 2026
Safety

Stress Testing Concept Erasure with Large Language Model Agents

DGX agent

arXiv:2607.17890v2 Announce Type: replace Abstract: Concept erasure aims to remove semantic concepts from a trained generative model and is increasingly important for responsible AI deployment. Howeve

safetyarxiv-cs-ai
23 Jul 2026
Safety

Toward Adaptable Multi-Agent Reinforcement Learning: An Assumption-Aware Review

DGX agent

arXiv:2507.10142v2 Announce Type: replace Abstract: Multi-Agent Reinforcement Learning (MARL) has achieved strong performance in simulated benchmarks, yet real deployments often violate the assumption

safetyarxiv-cs-ai
23 Jul 2026
Model Releases

Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents

DGX agent

arXiv:2511.18685v4 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) show promising results as decision-making engines for embodied agents operating in complex, physical enviro

model-releasesarxiv-cs-cv
16 Jul 2026
Model Releases

Quantum Circuit Vision: Cost-Aware Evaluation of Visual AI Agents for Quantum Code Generation

DGX agent

arXiv:2607.10057v1 Announce Type: cross Abstract: Can AI agents visually comprehend quantum circuit diagrams and generate verified executable code--and at what cost? We present Quantum Circuit Vision,

model-releasesarxiv-cs-cv
16 Jul 2026
Local Ai

A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism

DGX agent

arXiv:2607.12640v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards, and Group Relative Policy Optimization (GRPO) in particular, is now run routinely on a supervised checkp

local-aiarxiv-cs-ai
15 Jul 2026
Model Releases

Multi-Perspective Agentic Program Repair via Code Property Graphs and Temporal Execution Graphs

DGX agent

arXiv:2607.12605v1 Announce Type: cross Abstract: Large language models (LLMs) have improved automated program repair (APR), but two limitations remain. First, raw execution traces are often too large

model-releasesarxiv-cs-ai
15 Jul 2026
Safety

Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking

DGX agent

arXiv:2607.12085v1 Announce Type: new Abstract: Evaluating retail conversational agents requires methods beyond lexical-overlap metrics to assess intent alignment, factuality, helpfulness, clarity, to

safetyarxiv-cs-ai
15 Jul 2026
Agents

Agentic Neural Architecture Search

DGX agent

arXiv:2607.07984v1 Announce Type: new Abstract: Neural architecture search (NAS) methods have grown increasingly efficient, yet they remain bounded by manually engineered search spaces that require su

agentsarxiv-cs-ai
10 Jul 2026
Model Releases

CausalDS: Benchmarking Causal Reasoning in Data-Science Agents

DGX agent

arXiv:2607.08093v1 Announce Type: new Abstract: Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use. Yet the relevant b

model-releasesarxiv-cs-ai
10 Jul 2026
Model Releases

Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics

DGX agent

arXiv:2607.06820v1 Announce Type: new Abstract: Recent advances in AI for Mathematics have focused largely on autoformalization and theorem proving, leaving the role of Computer Algebra Systems (CAS)

model-releasesarxiv-cs-ai
9 Jul 2026
Safety

HiDVFS: Hierarchical Multi-Agent DVFS for Real-Time OpenMP DAG Workloads

DGX agent

arXiv:2601.06425v2 Announce Type: replace-cross Abstract: Leakage power in multicore embedded systems now rivals dynamic power, so DVFS schedulers must respect deadlines and thermal limits, not just a

safetyarxiv-cs-ai
9 Jul 2026
Safety

LLM-powered reasoning in agent-based modeling

DGX agent

arXiv:2607.06757v1 Announce Type: new Abstract: Agent-based modeling (ABM) has the capability to model millions of individuals and their interactions, which is useful for policy making. However, ABMs

safetyarxiv-cs-ai
9 Jul 2026
Agents

RIMRULE: Improving Tool-Using Language Agents via MDL-Guided Rule Learning

DGX agent

arXiv:2601.00086v3 Announce Type: replace Abstract: Large language models (LLMs) often struggle to use tools reliably in domain-specific settings, where APIs may be idiosyncratic, under-documented, or

agentsarxiv-cs-cl
9 Jul 2026
Safety

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

DGX agent

arXiv:2607.07508v1 Announce Type: cross Abstract: Reinforcement learning (RL) is becoming increasingly important for post-training large language models (LLMs). Previous RL pipelines for LLMs were mos

safetyarxiv-cs-ai
9 Jul 2026
Safety

Agentic AI for Commercial Insurance Underwriting with Adversarial Self-Critique

DGX agent

arXiv:2602.13213v2 Announce Type: replace Abstract: Commercial insurance underwriting is a labor-intensive process that requires manual review of extensive documentation to assess risk and determine p

safetyarxiv-cs-ai
8 Jul 2026
Agents

Agentic AI for IPoDWDM Network Lifecycle Automation: An MCP-Enabled Architecture

DGX agent

arXiv:2607.05958v1 Announce Type: cross Abstract: We present a distributed, vendor-agnostic multi-MCP architecture for SDN-based automation and autonomous control of multi-vendor, multi-layer IPoDWDM

agentsarxiv-cs-ai
8 Jul 2026
Safety

From Blueprint to Reality: Modeling and Applying Putnam's Social Capital Theory with LLM-based Multi-agent Simulations

DGX agent

arXiv:2607.06080v1 Announce Type: cross Abstract: Putnam's Social Capital Theory is a foundational framework for collective action and community prosperity. However, traditional empirical methods face

safetyarxiv-cs-ai
8 Jul 2026
Model Releases

IMR: Iterative Mode-World Weighted Regression for Multi-Agent Trajectory Prediction

DGX agent

arXiv:2607.05705v1 Announce Type: cross Abstract: Multi-agent motion prediction is essential for automated vehicles to understand the intentions of surrounding vehicles. However, previous prediction-b

model-releasesarxiv-cs-ai
8 Jul 2026
Safety

Intercepting an Agile Target with Net-Carrying Drones using Competitive Multi-Agent Reinforcement Learning

DGX agent

arXiv:2607.05939v1 Announce Type: new Abstract: This article presents a solution to intercept an agile drone by a team of agile drone carrying catching nets. We formulate the problem as a competitive

safetyarxiv-cs-ro
8 Jul 2026
Model Releases

Learning to Control LLM Agent Harnesses with Offline Reinforcement Learning

DGX agent

arXiv:2607.05458v1 Announce Type: cross Abstract: Large language model (LLM) agents are usually improved by changing prompts, models, or hand-written workflows, while the execution harness around the

model-releasesarxiv-cs-ai
8 Jul 2026
Safety

Lingering Authority: Revocable Resource-and-Effect Capabilities for Coding Agents

DGX agent

arXiv:2606.22504v1 Announce Type: cross Abstract: Coding agents often receive broad tool access for an entire task, even when a resource is needed only for one subgoal. We call this gap lingering auth

safetyarxiv-cs-ai
8 Jul 2026
Safety

TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training

DGX agent

arXiv:2607.05804v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student's own trajectories, offering a promising framework fo

safetyarxiv-cs-ai
8 Jul 2026
Model Releases

Agent Step Value: State-Transition Measurement with State-Grounded LLM Evaluators

DGX agent

arXiv:2607.04419v1 Announce Type: new Abstract: Most agent evaluations collapse a multi-step trace into a final answer, a success flag, or a trajectory-level score. These aggregates obscure the diagno

model-releasesarxiv-cs-ai
7 Jul 2026
Safety

Agentic Artificial Intelligence for Multistage Physics Experiments at a Large-Scale User Facility Particle Accelerator

DGX agent

arXiv:2509.17255v2 Announce Type: replace-cross Abstract: We present the first language-model-driven agentic artificial intelligence (AI) system to autonomously execute multi-stage physics experiments

safetyarxiv-cs-ai
7 Jul 2026
Model Releases

AgentLTL: A Trace-Verification Framework for Measuring, Enforcing, and Training Procedural Compliance in Tool-Using LLM Agents

DGX agent

arXiv:2607.02599v1 Announce Type: cross Abstract: Tool-using LLM agents are usually evaluated by final-answer correctness or LLM judges. Neither captures how an answer was produced. In safety-critical

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

Beyond Task Completion: A Verification-vs.-Conformance Gap in Tool-Evolving Agents

DGX agent

arXiv:2604.00392v2 Announce Type: replace-cross Abstract: Agents that synthesize their own tools ship a second artifact alongside each answer: a software library that future tasks reuse, compose, and

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning

DGX agent

arXiv:2601.15141v2 Announce Type: replace Abstract: Agentic Reinforcement Learning (RL) has empowered Large Language Models (LLMs) to utilize tools like Python interpreters for complex problem-solving

model-releasesarxiv-cs-lg
7 Jul 2026
Safety

GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation Tasks

DGX agent

arXiv:2607.05369v1 Announce Type: cross Abstract: For robots to work reliably in commercial and industrial applications, can recent advances in agentic coding systems combine interpretable robot progr

safetyarxiv-cs-ai
7 Jul 2026
Model Releases

Governed MCP: Kernel-Level Tool Governance for AI Agents via Logit-Based Safety Primitives

DGX agent

arXiv:2604.16870v2 Announce Type: replace-cross Abstract: AI agents increasingly call external tools (file system, network, APIs) through the Model Context Protocol (MCP). These tool calls are the age

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation

DGX agent

arXiv:2511.17384v2 Announce Type: replace-cross Abstract: While Visual Large Language Models (VLLMs) show great promise as embodied agents, they continue to face substantial challenges in spatial reas

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents

DGX agent

arXiv:2607.03968v1 Announce Type: cross Abstract: Large language models are increasingly deployed as IDE-integrated coding agents that decompose tasks, generate and edit files, run code, and refine ou

model-releasesarxiv-cs-ai
7 Jul 2026
Safety

Relational Multi-Agent Reinforcement Learning for Dynamic Pricing in High-Speed Railway Markets

DGX agent

arXiv:2607.05179v1 Announce Type: cross Abstract: In liberalised railway systems, operators must set prices dynamically in an environment with partial observability, as they retain private information

safetyarxiv-cs-ai
7 Jul 2026
Model Releases

The Remarkable Effectiveness of Providing AI Agents with Natural Language Tools: A Replication Study Validating NLT Performance Across 14 Models

DGX agent

arXiv:2607.03953v1 Announce Type: cross Abstract: This study independently replicates and extends the Natural Language Tools (NLT) framework of Johnson et al.~(2025), which questions the use of struct

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents

DGX agent

arXiv:2607.04686v1 Announce Type: cross Abstract: Tool calling is central to modern language model agents, but aggregate benchmark scores often hide where tool use fails. A model that never calls a ne

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

Towards Reliable Local Security Agents: Verifiable Post-Training for Linux Privilege Escalation

DGX agent

arXiv:2603.17673v2 Announce Type: replace-cross Abstract: LLM agents are becoming increasingly important in the security domain, but leading systems are often closed-source, cloud-based, hard to repro

model-releasesarxiv-cs-ai
7 Jul 2026
Safety

Transformer-Based Multi-Agent Reinforcement Learning for Networked Systems with Long-Range Interactions

DGX agent

arXiv:2511.13103v2 Announce Type: replace Abstract: Multi-agent reinforcement learning (MARL) has shown promise for large-scale network control, yet existing methods face two major limitations. First,

safetyarxiv-cs-lg
7 Jul 2026
← Previous
1…8687888990…236
Next →