AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,289 results
Model Releases

The Yes-Man Syndrome: Benchmarking Abstention in Embodied Robotic Agents

DGX agent

arXiv:2605.20544v1 Announce Type: cross Abstract: Vision-language models (VLMs) are used as high-level planners for embodied agents, translating natural language instructions and visual observations i

model-releasesarxiv-cs-cv
21 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

ContextFlow: Hierarchical Task-State Alignment for Long-Horizon Embodied Agents

DGX agent

arXiv:2605.19314v1 Announce Type: cross Abstract: Long-horizon embodied agents increasingly delegate navigation, search, approach, and manipulation to specialist executors. As these executors become s

safetyarxiv-cs-ai
20 May 2026
Model Releases

DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows

DGX agent

arXiv:2605.19099v1 Announce Type: new Abstract: We introduce DecisionBench, a benchmark substrate for emergent delegation in long-horizon agentic workflows. The substrate fixes a task suite (GAIA, tau

model-releasesarxiv-cs-ai
20 May 2026
Tutorials

Evaluating Memory Condensation Strategies for Coding Agents in Data-Driven Scientific Discovery

DGX agent

arXiv:2605.18854v1 Announce Type: new Abstract: Coding agents accumulate extensive context during long-running tasks, yet fixed context windows force practitioners to choose between truncation and tas

tutorialsarxiv-cs-lg
20 May 2026
Safety

Memory-Augmented Reinforcement Learning Agent for CAD Generation

DGX agent

arXiv:2605.19748v1 Announce Type: new Abstract: Automatic generation of computer-aided design (CAD) models is a core technology for enabling intelligence in advanced manufacturing. Existing generation

safetyarxiv-cs-ai
20 May 2026
Safety

Phase-Aware Mixture of Experts for Agentic Reinforcement Learning

DGX agent

arXiv:2602.17038v3 Announce Type: replace Abstract: Reinforcement learning (RL) has equipped LLM agents with a strong ability to solve complex tasks. However, existing RL methods normally use a single

safetyarxiv-cs-ai
20 May 2026
Model Releases

Rethinking How to Remember: Beyond Atomic Facts in Lifelong LLM Agent Memory

DGX agent

arXiv:2605.19952v1 Announce Type: new Abstract: To enable reliable long-term interaction, LLM agents require a memory system that can faithfully store, efficiently retrieve, and deeply reason over acc

model-releasesarxiv-cs-cl
20 May 2026
Safety

SimGym: A Framework for A/B Test Simulation in E-Commerce with Traffic-Grounded VLM Agents

DGX agent

arXiv:2605.19219v1 Announce Type: new Abstract: A/B testing remains the gold standard for evaluating modifications to e-commerce storefronts, yet it diverts traffic, requires weeks to reach statistica

safetyarxiv-cs-ai
20 May 2026
Model Releases

To Call or Not to Call: Diagnosing Intrinsic Over-Calling Bias in LLM Agents

DGX agent

arXiv:2605.18882v1 Announce Type: cross Abstract: LLM agents exhibit a consistent tendency to over-call, invoking tools even in situations where none is needed. On the When2Call benchmark, six models

model-releasesarxiv-cs-ai
20 May 2026
Local Ai

BLAgent: Agentic RAG for File-Level Bug Localization

DGX agent

arXiv:2605.17965v1 Announce Type: cross Abstract: Bug localization remains a key bottleneck in downstream software maintenance tasks, including root cause analysis, triage, and automated program repai

local-aiarxiv-cs-ai
19 May 2026
Model Releases

Causal Intervention-Based Memory Selection for Long-Horizon LLM Agents

DGX agent

arXiv:2605.17641v1 Announce Type: new Abstract: Long-horizon LLM agents rely on persistent memory to support interactions across sessions, yet existing memory systems often retrieve context using sema

model-releasesarxiv-cs-ai
19 May 2026
Local Ai

Enhancing Cloud Network Resilience via a Robust LLM-Empowered Multi-Agent Reinforcement Learning Framework

DGX agent

arXiv:2601.07122v2 Announce Type: replace-cross Abstract: While virtualization and resource pooling empower cloud networks with structural flexibility and elastic scalability, they inevitably expand t

local-aiarxiv-cs-ai
19 May 2026
Safety

Equilibrium Selection in Multi-Agent Policy Gradients via Opponent-Aware Basin Entry

DGX agent

arXiv:2605.18078v1 Announce Type: new Abstract: Multi-agent policy-gradient methods have been shown to converge locally near stable Nash equilibria. Local convergence, however, does not determine whic

safetyarxiv-cs-lg
19 May 2026
Model Releases

extsc{PrivScope}: Task-scoped Disclosure Control for Hybrid Agentic Systems

DGX agent

arXiv:2605.16630v1 Announce Type: cross Abstract: Hybrid local--cloud agents enrich user requests with context from persistent working state before delegating capability-intensive subtasks to a cloud

model-releasesarxiv-cs-ai
19 May 2026
Safety

Generation Navigator: A State-Aware Agentic Framework for Image Generation

DGX agent

arXiv:2605.17969v1 Announce Type: new Abstract: Despite rapid advances in text-to-image generation, faithfully realizing user intent remains challenging, often requiring manual multi-turn trial and er

safetyarxiv-cs-cv
19 May 2026
Model Releases

MAVEN A Multi-Agent Framework for Multicultural Text-to-Video Generation

DGX agent

arXiv:2605.16716v1 Announce Type: cross Abstract: Text-to-video (T2V) generation has rapidly progressed in visual fidelity, yet its ability to faithfully represent multiple cultures within a single pr

model-releasesarxiv-cs-ai
19 May 2026
Safety

Mitigating Conversational Inertia in Multi-Turn Agents

DGX agent

arXiv:2602.03664v3 Announce Type: replace Abstract: Large language models excel as few-shot learners when provided with appropriate demonstrations, yet this strength becomes problematic in multiturn a

safetyarxiv-cs-ai
19 May 2026
Local Ai

Robo-Cortex: A Self-Evolving Embodied Agent via Dual-Grain Cognitive Memory and Autonomous Knowledge Induction

DGX agent

arXiv:2605.18729v1 Announce Type: cross Abstract: The ability to navigate and interact with complex environments is central to real-world embodied agents, yet navigation in unseen environments remains

local-aiarxiv-cs-cv
19 May 2026
Model Releases

Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents

DGX agent

arXiv:2605.16986v1 Announce Type: cross Abstract: LLM agents benefit from reusable skills, yet test-time tasks often require guidance more specific than a static skill library can provide. We propose

model-releasesarxiv-cs-ai
19 May 2026
Research

SPIKE: An Adaptive Dual Controller Framework for Cost-Efficient Long-Horizon Game Agents

DGX agent

arXiv:2605.18636v1 Announce Type: new Abstract: Long-horizon multimodal agents in open-world games must stay goal-directed across many low-level interactions under tight token and latency budgets. Exi

researcharxiv-cs-cv
19 May 2026
Model Releases

TusoAI: Agentic Optimization for Scientific Methods

DGX agent

arXiv:2509.23986v2 Announce Type: replace Abstract: Scientific discovery is often slowed by the manual development of computational tools needed to analyze complex experimental data. Building such too

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

WEBSERV: A Full-Stack and RL-Ready Web Environment for Training Web Agents at Scale

DGX agent

arXiv:2510.16252v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) for web agents demands environments that are both effective for evaluation and efficient enough for large-scale on

model-releasesarxiv-cs-cl
19 May 2026
Safety

TopoEvo: A Topology-Aware Self-Evolving Multi-Agent Framework for Root Cause Analysis in Microservices

DGX agent

arXiv:2605.15611v1 Announce Type: new Abstract: Root cause analysis (RCA) in microservices is challenging due to (i) noisy and heterogeneous multimodal observability (metrics, logs, traces), (ii) casc

safetyarxiv-cs-ai
18 May 2026
Model Releases

AgenticEval: Toward Agentic and Self-Evolving Safety Evaluation of Large Language Models

DGX agent

arXiv:2509.26100v2 Announce Type: replace Abstract: The rapid integration of Large Language Models (LLMs) into high-stakes domains necessitates reliable safety and compliance evaluation. However, exis

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

EvolveMem:Self-Evolving Memory Architecture via AutoResearch for LLM Agents

DGX agent

arXiv:2605.13941v1 Announce Type: cross Abstract: Long-term memory is essential for LLM agents that operate across multiple sessions, yet existing memory systems treat retrieval infrastructure as fixe

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents

DGX agent

arXiv:2605.15104v1 Announce Type: new Abstract: Voice agents increasingly require reliable tool use from speech, whereas prominent tool-calling benchmarks remain text-based. We study whether verified

model-releasesarxiv-cs-cl
15 May 2026
Model Releases

Latency-Quality Routing for Functionally Equivalent Tools in LLM Agents

DGX agent

arXiv:2605.14241v1 Announce Type: new Abstract: Tool-augmented LLM agents increasingly access the same tool type through multiple functionally equivalent providers, such as web-search APIs, retrievers

model-releasesarxiv-cs-lg
15 May 2026
Model Releases

MemReranker: Reasoning-Aware Reranking for Agent Memory Retrieval

DGX agent

arXiv:2605.06132v2 Announce Type: replace Abstract: In agent memory systems, the reranking model serves as the critical bridge connecting user queries with long-term memory. Most systems adopt the 're

model-releasesarxiv-cs-cl
15 May 2026
Safety

Probabilistic Verification of Recurrent Neural Networks for Single and Multi-Agent Reinforcement Learning

DGX agent

arXiv:2605.14758v1 Announce Type: new Abstract: History-dependent policies induced by recurrent neural networks (RNNs) rely on latent hidden state dynamics, making verification in partially observable

safetyarxiv-cs-ai
15 May 2026
Model Releases

Reinforcement Learning for Tool-Calling Agents in Fast Healthcare Interoperability Resources (FHIR)

DGX agent

arXiv:2605.14126v1 Announce Type: cross Abstract: Fast Healthcare Interoperability Resources (FHIR) is the dominant standard for interoperable exchange of healthcare data. In FHIR, electronic health r

model-releasesarxiv-cs-ai
15 May 2026
Safety

Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy

DGX agent

arXiv:2605.14558v1 Announce Type: cross Abstract: Agentic reinforcement learning trains large language models using multi-turn trajectories that interleave long reasoning traces with short environment

safetyarxiv-cs-ai
15 May 2026
Local Ai

Towards In-Depth Root Cause Localization for Microservices with Multi-Agent Recursion-of-Thought

DGX agent

arXiv:2605.14866v1 Announce Type: cross Abstract: As modern microservice systems grow increasingly complex due to dynamic interactions and evolving runtime environments, they experience failures with

local-aiarxiv-cs-ai
15 May 2026
Local Ai

Agentic Interpretation: Lattice-Structured Evidence for LLM-Based Program Analysis

DGX agent

arXiv:2605.12694v1 Announce Type: cross Abstract: Large language models can consult information that fixed static analyzers cannot, such as documentation, current security advisories, version-specific

local-aiarxiv-cs-ai
14 May 2026
Safety

Data Agent: Learning to Select Data via End-to-End Dynamic Optimization

DGX agent

arXiv:2603.07433v2 Announce Type: replace-cross Abstract: Dynamic Data selection aims to accelerate training by prioritizing informative samples during online training. However, existing methods typic

safetyarxiv-cs-cv
14 May 2026
Safety

Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy

DGX agent

arXiv:2605.12991v1 Announce Type: cross Abstract: LLM-based multi-agent pipelines flip from correct to incorrect answers under simulated peer disagreement at rates we term yield, a vulnerability widel

safetyarxiv-cs-ai
14 May 2026
Safety

VideoSEAL: Mitigating Evidence Misalignment in Agentic Long Video Understanding by Decoupling Answer Authority

DGX agent

arXiv:2605.12571v1 Announce Type: cross Abstract: Long video question answering requires locating sparse, time-scattered visual evidence within highly redundant content. Although current MLLMs perform

safetyarxiv-cs-ai
14 May 2026
Safety

Adaptive TD-Lambda for Cooperative Multi-agent Reinforcement Learning

DGX agent

arXiv:2605.11880v1 Announce Type: new Abstract: TD(lambda) in value-based MARL algorithms or the Temporal Difference critic learning in Actor-Critic-based (AC-based) algorithms synergistically integra

safetyarxiv-cs-lg
13 May 2026
Model Releases

Agent-Based Post-Hoc Correction of Agricultural Yield Forecasts

DGX agent

arXiv:2605.12375v1 Announce Type: new Abstract: Accurate crop yield forecasting in commercial soft fruit production is constrained by the data available in typical commercial farm records, which lack

model-releasesarxiv-cs-lg
13 May 2026
Model Releases

Courtroom-Style Multi-Agent Debate with Progressive RAG and Role-Switching for Controversial Claim Verification

DGX agent

arXiv:2603.28488v2 Announce Type: replace Abstract: Large language models (LLMs) remain unreliable for high-stakes claim verification due to hallucinations and shallow reasoning. While retrieval-augme

model-releasesarxiv-cs-cl
13 May 2026
Safety

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction

DGX agent

arXiv:2605.12070v1 Announce Type: new Abstract: Asynchronous reinforcement learning improves rollout throughput for large language model agents by decoupling sample generation from policy optimization

safetyarxiv-cs-lg
13 May 2026
Model Releases

PRISM: Pareto-Efficient Retrieval over Intent-Aware Structured Memory for Long-Horizon Agents

DGX agent

arXiv:2605.12260v1 Announce Type: new Abstract: Long-horizon language agents accumulate conversation history far faster than any fixed context window can hold, making memory management critical to bot

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

A Communication-Theoretic Framework for LLM Agents: Cost-Aware Adaptive Reliability

DGX agent

arXiv:2605.09121v1 Announce Type: cross Abstract: Agents built on large language models (LLMs) rely on a range of reliability techniques, including retry, majority voting, and self-consistency, that h

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

AnomalyClaw: A Universal Visual Anomaly Detection Agent via Tool-Grounded Refutation

DGX agent

arXiv:2605.10397v1 Announce Type: cross Abstract: Visual anomaly detection (VAD) is crucial in many real-world fields, such as industrial inspection, medical imaging, infrastructure monitoring, and re

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

AssemPlanner: A Multi-Agent Based Task Planning Framework for Flexible Assembly System

DGX agent

arXiv:2605.08831v1 Announce Type: new Abstract: In flexible assembly systems, existing task planning methods require a time-consuming configuration process by multiple experts to establish a productio

model-releasesarxiv-cs-ro
12 May 2026
Model Releases

CIVeX: Causal Intervention Verification for Language Agents

DGX agent

arXiv:2605.09168v1 Announce Type: new Abstract: A valid tool call is not necessarily a valid intervention. Tool-using language agents are guarded by schema validators, policy filters, provenance check

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Combining Mechanical and Agentic Specification Inference for Move

DGX agent

arXiv:2605.10005v1 Announce Type: cross Abstract: In this paper, we describe early work on a specification inference tool for the Move Prover that combines a weakest-precondition (WP) analysis over Mo

model-releasesarxiv-cs-ai
12 May 2026
Safety

Conformity Generates Collective Misalignment in AI Agents Societies

DGX agent

arXiv:2605.10721v1 Announce Type: cross Abstract: Artificial intelligence safety research focuses on aligning individual language models with human values, yet deployed AI systems increasingly operate

safetyarxiv-cs-cl
12 May 2026
Research

Defense effectiveness across architectural layers: a mechanistic evaluation of persistent memory attacks on stateful LLM agents

DGX agent

arXiv:2605.08442v1 Announce Type: cross Abstract: Persistent memory attacks against LLM agents achieve high attack success rates against open-source models. In these attacks, malicious instructions in

researcharxiv-cs-ai
12 May 2026
← Previous
1…101102103104105…236
Next →