AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,289 results
Local Ai

RIZZ: Routing Interactions to Near Zero-Interference Zones for Continual Adaptation of Black-Box Agents

DGX agent

arXiv:2606.20638v1 Announce Type: cross Abstract: Large language models are increasingly deployed as long-lived agents that must adapt across users, tasks, domains, modalities, and feedback regimes wi

local-aiarxiv-cs-lg
23 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

RoboLineage: Agent-Native Data Lifecycle Governance Across Robot Policy Iterations

DGX agent

arXiv:2606.22142v1 Announce Type: new Abstract: We present RoboLineage, an agent-native data lifecycle governance system for robot policy iteration. Modern robot policies improve through repeated data

safetyarxiv-cs-ro
23 Jun 2026
Safety

Sovereign Execution Broker: Enforcing Certificate-Bound Authority in Agentic Control Planes

DGX agent

arXiv:2606.20520v2 Announce Type: replace-cross Abstract: Autonomous agents are increasingly connected to cloud, deployment, and data-control workflows, but production mutation authority should not re

safetyarxiv-cs-lg
23 Jun 2026
Model Releases

When Web Agents Finish but Still Fail: Reproducible Triggers and Trace Diagnostics for Parallel Web Exploration

DGX agent

arXiv:2606.20724v1 Announce Type: cross Abstract: Long-horizon web agents often fail in ways hidden by final-answer evaluation: they may visit useful pages, produce a well-formed answer, and terminate

model-releasesarxiv-cs-lg
23 Jun 2026
Model Releases

A Lightweight Multi-Agent Framework for Automated Concrete Barrier Design

DGX agent

arXiv:2606.12040v1 Announce Type: new Abstract: The design of reinforced concrete highway barriers is a safety-critical process that requires strict compliance with regulatory provisions such as the A

model-releasesarxiv-cs-ai
11 Jun 2026
Safety

AerialClaw: An Open-Source Framework for LLM-Driven Autonomous Aerial Agents

DGX agent

arXiv:2606.12142v1 Announce Type: cross Abstract: Unmanned aerial vehicles (UAVs) are increasingly used in inspection, search and rescue, environmental monitoring, and emergency response. However, mos

safetyarxiv-cs-cv
11 Jun 2026
Research

FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents

DGX agent

arXiv:2606.12087v1 Announce Type: new Abstract: Training deep search agents requires verifiable questions whose answers remain unavailable until sufficient evidence has been acquired through search. E

researcharxiv-cs-cl
11 Jun 2026
Safety

Human-Enhanced Loop Modeling (HELM): Agent-Based Finite Element Modeling of Concrete Bridge Barriers

DGX agent

arXiv:2606.12025v1 Announce Type: new Abstract: Finite element (FE) modeling of safety-critical infrastructure such as bridge barriers requires high-fidelity nonlinear dynamic analysis, yet the curren

safetyarxiv-cs-ai
11 Jun 2026
Safety

IAPO: Input Attribution-Aware Policy Optimization for Tool Use in Small Multimodal Agents

DGX agent

arXiv:2606.11652v1 Announce Type: new Abstract: This paper investigates reinforcement learning (RL) methods for improving tool-calling capabilities in multimodal small language model (SLM) agents. Whi

safetyarxiv-cs-lg
11 Jun 2026
Model Releases

NightFeats @ MMU-RAGent NeurIPS 2025: A Context-Optimized Multi-Agent RAG System for the Text-to-Text Track

DGX agent

arXiv:2606.11199v1 Announce Type: cross Abstract: We present NightFeats, a structured multi-agent retrieval-augmented generation (RAG) system submitted to the MMU-RAGent competition at NeurIPS 2025, w

model-releasesarxiv-cs-ai
11 Jun 2026
Safety

Sovereign Assurance Boundary: Certificate-Bound Admission for Agentic Infrastructure

DGX agent

arXiv:2606.11632v1 Announce Type: cross Abstract: Agentic infrastructure introduces a critical control-plane authorization problem: non-deterministic reasoning systems can propose high-stakes mutation

safetyarxiv-cs-ai
11 Jun 2026
Safety

3SPO: State-Score-Supervised Policy Optimization for LLM Agents

DGX agent

arXiv:2606.09961v1 Announce Type: cross Abstract: Training large language models (LLMs) as autonomous agents via reinforcement learning (RL) has enabled frontier models to achieve superhuman performan

safetyarxiv-cs-ai
10 Jun 2026
Model Releases

ASA: Backbone-Training-Free Representation Engineering for Tool-Calling Agents

DGX agent

arXiv:2602.04935v3 Announce Type: replace-cross Abstract: Adapting LLM agents to domain-specific tool calling remains notably brittle under evolving interfaces. Prompt and schema engineering is easy t

model-releasesarxiv-cs-ai
10 Jun 2026
Safety

AsyncWebRL: Efficient Multi-Step RL for Visual Web Agents

DGX agent

arXiv:2606.05597v2 Announce Type: replace Abstract: Training vision-language web agents with multi-step RL is compute-intensive, with two dominant forms of inefficiency: idle GPUs in synchronous RL, a

safetyarxiv-cs-lg
10 Jun 2026
Model Releases

Can Multi-Agent LLMs Identify Their Peers? Stylometric Fingerprinting in Role-Constrained Political Analysis

DGX agent

arXiv:2606.09854v1 Announce Type: cross Abstract: Multi-agent large language model (LLM) pipelines for political statement analysis are vulnerable to peer-preservation bias: models tend to protect pee

model-releasesarxiv-cs-ai
10 Jun 2026
Safety

Decoupling Thought from Speech: Knowledge-Grounded Counterfactual Reasoning for Resilient Multi-Agent Argumentation

DGX agent

arXiv:2606.10475v1 Announce Type: cross Abstract: Multi-agent debate frameworks have been shown to improve large language model performance in convergent tasks, but they are currently optimized in a w

safetyarxiv-cs-ai
10 Jun 2026
Model Releases

IntentKV: Cross-Turn Intent-Aware KV Cache Pruning for Agent Inference

DGX agent

arXiv:2606.09916v1 Announce Type: cross Abstract: Multi-turn LLM agents fan short queries into long trajectories of tool calls, search results, and intermediate reasoning. Both KV memory and KV read b

model-releasesarxiv-cs-ai
10 Jun 2026
Local Ai

Learning What to Remember: Observability-Safe Memory Retention via Constrained Optimization for Long-Horizon Language Agents

DGX agent

arXiv:2606.10616v1 Announce Type: new Abstract: Long-horizon language agents accumulate observations, reasoning traces, and retrieved facts that exceed their finite context windows, making memory rete

local-aiarxiv-cs-ai
10 Jun 2026
Model Releases

RedAct: Redacting Agent Capability Traces for Procedural Skill Protection

DGX agent

arXiv:2606.10813v1 Announce Type: cross Abstract: Users rely on execution traces to observe agent behavior, diagnose failures, and ensure accountability. These traces contain rich procedural detail, i

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

SkillResolve-Bench: Measuring and Resolving Same-Capability Ambiguity in Agent Skill Retrieval

DGX agent

arXiv:2606.10388v1 Announce Type: cross Abstract: Agent skill libraries are becoming routable software assets: a retrieved skill can contribute instructions, scripts, resource bindings, and execution

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

AliyunConsoleAgent: Training Web Agents in Real-World Cloud Environments via Distillation and Reinforcement Learning

DGX agent

arXiv:2606.09447v1 Announce Type: new Abstract: We present AliyunConsoleAgent, a web agent framework for automated documentation verification in real-world cloud consoles. Major cloud platforms encomp

model-releasesarxiv-cs-ai
9 Jun 2026
Safety

Crayotter: Traceable Multi-Agent Workflows for Long-Form Video Editing

DGX agent

arXiv:2606.07636v1 Announce Type: new Abstract: Editing a long-form video from heterogeneous footage requires more than selecting clips: an agent must preserve narrative intent across material prepara

safetyarxiv-cs-cv
9 Jun 2026
Safety

Does Persona Make LLMs K-pop Fans? A Pilot Study of LLM-Based Online Concert Audience Agents

DGX agent

arXiv:2606.07837v1 Announce Type: cross Abstract: A concert is a collective experience, but recorded performance videos are typically watched alone, stripping away the shared audience presence that ma

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

GRPO Does Not Close the Multi-Agent Coordination Gap

DGX agent

arXiv:2606.07845v1 Announce Type: cross Abstract: We measure how well current large language models coordinate as multiple agents sharing a common resource, using the dining philosophers problem as a

model-releasesarxiv-cs-lg
9 Jun 2026
Research

IRAM-Omega-Q: A Computational Framework for Uncertainty Regulation in Adaptive Agents

DGX agent

arXiv:2603.16020v2 Announce Type: replace Abstract: Adaptive agents operating under uncertainty must do more than optimize task outputs: they must maintain a workable internal state under noise, pertu

researcharxiv-cs-ai
9 Jun 2026
Model Releases

Memory Beyond Recall: A Dual-Process Cognitive Memory System for Self-Evolving LLM Agents

DGX agent

arXiv:2606.09483v1 Announce Type: cross Abstract: Long-term memory for an LLM agent is more than retrieving the right passage at the right time. Current memory systems collapse belief revision, causal

model-releasesarxiv-cs-ai
9 Jun 2026
Safety

Payoff scaling shapes cooperation in LLM agents across languages

DGX agent

arXiv:2601.19082v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as autonomous agents that negotiate, coordinate, and act on behalf of users. Whether they coo

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

SceneConductor: 3D Scene Generation from Single Image with Multi-Agent Orchestration

DGX agent

arXiv:2606.08402v1 Announce Type: cross Abstract: Generating complete 3D scenes from a single image requires inferring globally consistent geometry, object relationships, and environmental context fro

model-releasesarxiv-cs-ai
9 Jun 2026
Safety

SciTrace: Trajectory-Aware Safety Reasoning for Scientific Discovery Agents

DGX agent

arXiv:2606.08234v1 Announce Type: new Abstract: LLM-based scientific agents have shown strong capacity for autonomous research, yet their safety layers remain structurally divorced from core reasoning

safetyarxiv-cs-ai
9 Jun 2026
Safety

SecureClaw: Clawing Back Control of LLM Agents

DGX agent

arXiv:2606.09549v1 Announce Type: cross Abstract: Tool-using large language model (LLM) agents face two distinct security failures: unauthorized external actions and exposure of sensitive plaintext in

safetyarxiv-cs-ai
9 Jun 2026
Safety

Self-Evolving Scientific Agent Discovers Generalizable Physically-Reasoned Fluid Control

DGX agent

arXiv:2606.08405v1 Announce Type: new Abstract: While data-intensive deep reinforcement learning can optimize complex control policies, scientific discovery in physical systems fundamentally requires

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks

DGX agent

arXiv:2606.09669v1 Announce Type: new Abstract: Spatial reasoning is a foundational capability for multimodal large language models (MLLMs) to perceive and operate within the physical world. However,

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

VisualLeakBench: Reproducible Action-Boundary Propagation Failures in Vision-Language Agents

DGX agent

arXiv:2606.07595v1 Announce Type: cross Abstract: Vision-language agents increasingly consume screenshots, documents, and user interfaces before writing to memory, sending messages, or invoking extern

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents

DGX agent

arXiv:2602.08235v2 Announce Type: replace-cross Abstract: Although computer-use agents (CUAs) hold significant potential to automate increasingly complex OS workflows, they can demonstrate unsafe unin

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

Audio-Visual World Models: Grounding Multisensory Imagination for Embodied Agents

DGX agent

arXiv:2512.00883v3 Announce Type: replace-cross Abstract: World models simulate environmental dynamics to enable agents to plan and reason about future states. While existing approaches have primarily

model-releasesarxiv-cs-cv
8 Jun 2026
Agents

Feasible Action Space Reduction for Quantifying Causal Responsibility in Continuous Spatial Interactions

DGX agent

arXiv:2505.17739v2 Announce Type: replace-cross Abstract: Understanding the causal influence of one agent on another agent is crucial for safely deploying artificially intelligent systems such as auto

agentsarxiv-cs-ro
8 Jun 2026
Model Releases

Hierarchical Certified Semantic Commitment for Byzantine-Resilient LLM-Agent Collaboration

DGX agent

arXiv:2606.07316v1 Announce Type: cross Abstract: Byzantine collaboration among large-language-model agents requires a finality-control primitive: given delivered stochastic, structured natural-langua

model-releasesarxiv-cs-ai
8 Jun 2026
Safety

Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates

DGX agent

arXiv:2601.18510v2 Announce Type: replace-cross Abstract: While Large Language Model (LLM) agents excel at general tasks, they inherently struggle with continual adaptation due to the frozen weights a

safetyarxiv-cs-ai
8 Jun 2026
Model Releases

M^3Exam: Benchmarking Multimodal Memory for Realistic User-Agent Interactions

DGX agent

arXiv:2606.07402v1 Announce Type: new Abstract: Language agents are increasingly deployed over accumulating multimodal information, yet existing benchmarks assume a human-human form with sparse visual

model-releasesarxiv-cs-cl
8 Jun 2026
Model Releases

MacArena: Benchmarking Computer Use Agents on an Online macOS Environment

DGX agent

arXiv:2606.06560v1 Announce Type: cross Abstract: Computer-use agents (CUAs) operate graphical user interfaces (GUIs) through vision and control primitives, and their capabilities have advanced rapidl

model-releasesarxiv-cs-ai
8 Jun 2026
Local Ai

StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents

DGX agent

arXiv:2606.07027v1 Announce Type: new Abstract: Reinforcement Learning (RL) has become a promising approach for improving GUI Agents in long-horizon, stochastic digital environments, but trajectory-le

local-aiarxiv-cs-ai
8 Jun 2026
Model Releases

DPBench: Structural Determinants of Multi-Agent LLM Coordination Under Simultaneous Resource Contention

DGX agent

arXiv:2602.13255v2 Announce Type: replace Abstract: We present DPBench, a benchmark for evaluating coordination in multi-agent systems built from large language models. Existing benchmarks measure tas

model-releasesarxiv-cs-ai
6 Jun 2026
Safety

Policy-Conditioned Counterfactual Credit for Verifiable Reinforcement Learning of Long-Horizon Language Agents

DGX agent

arXiv:2606.05263v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards improves reasoning and tool use, yet long-horizon language agents still learn unsupported evidence chai

safetyarxiv-cs-ai
6 Jun 2026
Model Releases

Domain-Conditioned Safety in Frontier Computer-Using Agents: A 793-Episode Browser Benchmark, a Coding-Domain Cross-Reference, and a Reproducibility Audit of Recent Red-Teaming

DGX agent

arXiv:2606.05233v1 Announce Type: cross Abstract: Recent computer-using-agent (CUA) red-teaming papers report prompt-injection attack success rates (ASR) of 42-98%, but these headline numbers cluster

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents

DGX agent

arXiv:2606.06087v1 Announce Type: new Abstract: Agent systems increasingly use textual skills to encode reusable task procedures, but injecting these skills into the prompt at every step incurs substa

model-releasesarxiv-cs-cl
5 Jun 2026
Local Ai

VASO: Formally Verifiable Self-Evolving Skills for Physical AI Agents

DGX agent

arXiv:2606.05395v1 Announce Type: new Abstract: Reusable robot skills are becoming the basic units through which embodied agents turn open-ended instructions into long-horizon physical behavior. We ar

local-aiarxiv-cs-ro
5 Jun 2026
Research

ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents

DGX agent

arXiv:2407.03884v4 Announce Type: replace-cross Abstract: Dialogue agents powered by Large Language Models (LLMs) show superior performance in various tasks. Despite the better user understanding and

researcharxiv-cs-ai
4 Jun 2026
Safety

Multi-Agent Next-Best-View Optimization for Risk-Averse Planning

DGX agent

arXiv:2606.04158v1 Announce Type: new Abstract: Multi-agent Next-Best-View (NBV) selection for safe path planning in uncertain and unknown environments requires informative, safety-aware, and efficien

safetyarxiv-cs-ro
4 Jun 2026
← Previous
1…8889909192…236
Next →