AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,289 results
Safety

PRO-CUA: Process-Reward Optimization for Computer Use Agents

DGX agent

arXiv:2605.29119v1 Announce Type: new Abstract: Computer use agents (CUAs) have shown strong potential for automating complex digital workflows, yet their training remains constrained by costly live e

safetyarxiv-cs-ai
29 May 2026
Safety
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Representation Signatures and Risk-Feedback Alignment in LLM Trading Agents

DGX agent

arXiv:2605.28850v1 Announce Type: new Abstract: We study behavioral alignment and representation dynamics of large language model (LLM) agents in financial decision environments. Using TradeArena, an

safetyarxiv-cs-lg
29 May 2026
Model Releases

SAAS: Self-Aware Reinforcement Learning for Over-Search Mitigation in Agentic Search

DGX agent

arXiv:2605.29796v1 Announce Type: new Abstract: Agentic search enables LLMs to solve complex multi-hop questions through iterative reasoning and external search. Despite the effectiveness, these syste

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

SkillsInjector: Dynamic Skill Context Construction for LLM Agents

DGX agent

arXiv:2605.29794v1 Announce Type: new Abstract: LLM agents now draw on growing skill libraries to handle complex tasks. However, injecting more skills does not always improve task completion and can e

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems

DGX agent

arXiv:2602.15382v2 Announce Type: replace Abstract: Multi-Agent Systems (MAS) powered by Large Language Models have unlocked advanced collaborative reasoning, yet they remain bottlenecked by discrete

model-releasesarxiv-cs-cl
29 May 2026
Local Ai

Unifying Temporal and Structural Credit Assignment in LLM-Based Multi-Agent Prompt Optimization

DGX agent

arXiv:2605.30227v1 Announce Type: cross Abstract: While Multi-Agent Systems (MAS) empower Large Language Models to tackle complex reasoning tasks through collaborative interaction, optimizing their dy

local-aiarxiv-cs-ai
29 May 2026
Model Releases

Agentic Active Omni-Modal Perception for Multi-Hop Audio-Visual Reasoning

DGX agent

arXiv:2605.28192v1 Announce Type: new Abstract: Multi-hop audio-visual reasoning remains challenging for Omni-LLMs, as relevant evidence is often sparse, temporally dispersed, and distributed across b

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

AndroidDaily: A Verifiable Benchmark for Mobile GUI Agents on Real-World Closed-Source Applications

DGX agent

arXiv:2605.27761v1 Announce Type: new Abstract: The rapid development of GUI foundation models and mobile GUI agents has spurred numerous evaluation benchmarks, yet most rely on simulated environments

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

Do Agents Think Deeper? A Mechanistic Investigation of Layer-Wise Dynamics in Sequential Planning

DGX agent

arXiv:2605.27935v1 Announce Type: new Abstract: Recent mechanistic studies suggest that large language models (LLMs) may utilize their depth inefficiently in standard single-turn tasks. Whether this s

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Dr-CiK: A Testbed for Foresight-Driven Agents

DGX agent

arXiv:2605.27904v1 Announce Type: new Abstract: Time series forecasting in real-world settings often depends not only on historical observations, but also on external context that must be actively dis

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

DynaSchedBench: Calibrated Dynamic Scheduling Benchmarks and Observability Paradox in LLM-based Scheduling Agents

DGX agent

arXiv:2605.27566v1 Announce Type: new Abstract: Progress in neural combinatorial optimization for Dynamic Flexible Job Shop Scheduling Problem (DFJSP) is currently hindered by a methodological tension

model-releasesarxiv-cs-ai
28 May 2026
Agents

Fine-Tuning Vision-Language Models for Understanding Current Damage and Scoring Priority with Quality Guard Agent

DGX agent

arXiv:2605.27452v1 Announce Type: new Abstract: Bridge inspection in Japan requires mandatory visual assessments every five years, yet qualitative damage ratings (levels a-e) assigned by different eng

agentsarxiv-cs-cv
28 May 2026
Model Releases

MaskClaw: Edge-Side Personalized Privacy Arbitration for GUI Agents with Behavior-Driven Skill Evolution

DGX agent

arXiv:2605.28646v1 Announce Type: cross Abstract: GUI agents rely on screenshots to infer intent and operate across applications, but these screenshots often contain private messages, medical records,

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Tool Forge: A Validation-Carrying Toolchain for Governed Agentic Execution

DGX agent

arXiv:2605.28000v1 Announce Type: cross Abstract: Large language model agents are increasingly expected to perform operational work: calling APIs, manipulating files, assembling workflows, and acting

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Anchor: Mitigating Artifact Drift in Agent Benchmark Generation

DGX agent

arXiv:2605.26321v1 Announce Type: new Abstract: AI agents are beginning to complete valuable, long-horizon business operations tasks, but training and evaluation environments for enterprise work still

model-releasesarxiv-cs-ai
27 May 2026
Agents

APEX-Searcher: Refining Credit Assignment with Subgoaling for Agentic Retrieval-Augmented Generation

DGX agent

arXiv:2603.13853v3 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) connects large language models (LLMs) to external knowledge, but single-round retrieval is often insuffic

agentsarxiv-cs-ai
27 May 2026
Safety

EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation

DGX agent

arXiv:2605.26785v1 Announce Type: cross Abstract: Post-trained LLMs are often optimized to align responses with human preferences, making them safe, polite, and conversationally appropriate. In advers

safetyarxiv-cs-ai
27 May 2026
Model Releases

Shopping Companion: A Memory-Augmented LLM Agent for Real-World E-Commerce Tasks

DGX agent

arXiv:2603.14864v2 Announce Type: replace Abstract: In e-commerce, LLM agents show promise for shopping tasks such as recommendations, budget management, and bundle deals, where accurately capturing u

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Tool-Schema Compression Enables Agentic RAG Under Constrained Context Budgets

DGX agent

arXiv:2605.26165v1 Announce Type: cross Abstract: Agentic RAG systems that equip language models with dozens to hundreds of tool definitions face a critical resource conflict: tool schemas consume the

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

TowerMind: A Tower Defence Game Learning Environment and Benchmark for LLM as Agents

DGX agent

arXiv:2601.05899v2 Announce Type: replace Abstract: Recent breakthroughs in Large Language Models (LLMs) have positioned them as a promising paradigm for agents, with long-term planning and decision-m

model-releasesarxiv-cs-ai
27 May 2026
Safety

Agent-Facing Information Design in LLM Tool Registries

DGX agent

arXiv:2605.23916v1 Announce Type: cross Abstract: LLM tool registries function as unregulated advertising platforms: providers write free-text descriptions that agents use for selection, yet no measur

safetyarxiv-cs-ai
26 May 2026
Research

DRIVE: Modeling Skills at the Reasoning and Interaction Levels for Web Agents under Continual Learning

DGX agent

arXiv:2605.23939v1 Announce Type: new Abstract: Web agents require both high-level reasoning (for task decomposition) and low-level interactions (for page elements manipulation) to conduct different t

researcharxiv-cs-ai
26 May 2026
Safety

Dynamic Dual-Granularity Skill Bank for Agentic RL

DGX agent

arXiv:2603.28716v2 Announce Type: replace Abstract: Agentic RL can benefit substantially from reusable experience, yet existing skill-based methods mainly extract trajectory-level guidance and often l

safetyarxiv-cs-ai
26 May 2026
Safety

ECHO: Terminal Agents Learn World Models for Free

DGX agent

arXiv:2605.24517v1 Announce Type: cross Abstract: CLI agents are the closest thing language models have to an embodied setting: the model emits commands, the terminal executes them, and the returned s

safetyarxiv-cs-cl
26 May 2026
Model Releases

Memory-Induced Tool-Drift in LLM Agents

DGX agent

arXiv:2605.24941v1 Announce Type: cross Abstract: Modern LLM agents combine long-term memory for personalization with tool-calling interfaces for taking actions in the world -- a combination underpinn

model-releasesarxiv-cs-lg
26 May 2026
Research

Mitigating Provenance-Role Collapse in Long-Term Agents via Typed Memory Representation

DGX agent

arXiv:2605.25869v1 Announce Type: new Abstract: Long-term memory is essential for persistent LLM agents, yet prevailing architectures store historical interactions as unstructured, flat text. This unc

researcharxiv-cs-cl
26 May 2026
Agents

MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research

DGX agent

arXiv:2605.26114v1 Announce Type: new Abstract: We present MobileGym, a browser-hosted, lightweight, fully controllable environment for everyday mobile use, targeting interaction fidelity without repl

agentsarxiv-cs-ai
26 May 2026
Model Releases

VeriTrace: Evolving Mental Models for Deep Research Agents

DGX agent

arXiv:2605.26081v1 Announce Type: new Abstract: Deep research agents face vast, interdependent, and pervasively uncertain information. Existing systems explore what evolving intermediate representatio

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

When Do LLM Agents Treat Surface Noise Differently from Semantic Noise? A 68-Cell Measurement Study with a Held-Out Trace-Level Validation

DGX agent

arXiv:2605.25981v1 Announce Type: new Abstract: We document an empirical phenomenon in chain-of-thought and ReAct agents driven by ten large language models from seven architecture families: meaning-b

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Ax-Prover: A Deep Reasoning Agentic Framework for Theorem Proving in Mathematics and Quantum Physics

DGX agent

arXiv:2510.12787v4 Announce Type: replace Abstract: We present Ax-Prover, a multi-agent system for automated theorem proving in Lean that can solve problems across diverse scientific domains and opera

model-releasesarxiv-cs-ai
25 May 2026
Safety

Goal-Conditioned Agents that Learn Everything All at Once

DGX agent

arXiv:2605.23551v1 Announce Type: cross Abstract: A goal-conditioned reinforcement learning agent exploring an environment will see a wealth of information throughout a trajectory, most of which is di

safetyarxiv-cs-ai
25 May 2026
Local Ai

HawkesLLM: Semantic Uncertainty Propagation in Agentic Text Simulation

DGX agent

arXiv:2605.23043v1 Announce Type: new Abstract: Agentic text-simulation systems write in sequence, with each item becoming possible context for later steps. That makes uncertainty path-dependent: an e

local-aiarxiv-cs-cl
25 May 2026
Model Releases

LLM-driven design of physics-constrained constitutive models: two agents are better than one

DGX agent

arXiv:2605.23754v1 Announce Type: new Abstract: Developing constitutive models that capture how materials deform under load traditionally requires years of specialized expertise in continuum mechanics

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

What Training Data Teaches RL Memory Agents: An Empirical Study of Curriculum Effects in Memory-Augmented QA

DGX agent

arXiv:2605.23067v1 Announce Type: new Abstract: Reinforcement learning (RL) has emerged as a viable recipe for training LLM agents to reason over external memory banks in multi-session dialogue. Exist

model-releasesarxiv-cs-cl
25 May 2026
Research

WMAttack: Automated Attack Search for Adversarial Evaluation of World-Model Agents

DGX agent

arXiv:2605.23220v1 Announce Type: new Abstract: Despite the growing use of world models as decision-making agents, their adversarial robustness remains underexplored due to the lack of dedicated autom

researcharxiv-cs-lg
25 May 2026
Model Releases

Blind Spots in the Guard: How Domain-Camouflaged Injection Attacks Evade Detection in Multi-Agent LLM Systems

DGX agent

arXiv:2605.22001v1 Announce Type: cross Abstract: Injection detectors deployed to protect LLM agents are calibrated on static, template-based payloads that announce themselves as override directives.

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Diverse Yet Consistent: Context-Guided Diffusion with Energy-Based Joint Refinement for Multi-Agent Motion Prediction

DGX agent

arXiv:2605.22017v1 Announce Type: new Abstract: Deepgenerative models havebecomeapromisingapproach for human motion prediction due to their ability to capture multimodal distributions and represent di

model-releasesarxiv-cs-cv
22 May 2026
Agents

SpecHop: Continuous Speculation for Accelerating Multi-Hop Retrieval Agents

DGX agent

arXiv:2605.21965v1 Announce Type: new Abstract: Large language models increasingly use external tools such as web search and document retrieval to solve information-intensive tasks. However, multi-hop

agentsarxiv-cs-cl
22 May 2026
Local Ai

VBFDD-Agent for Electric Vehicle Battery Fault Detection and Diagnosis: Descriptive Text Modeling of Battery Digital Signals

DGX agent

arXiv:2605.20742v1 Announce Type: new Abstract: With the rapid proliferation of electric vehicles, the safety and reliability of lithium-ion batteries have become critical concerns. Effective anomaly

local-aiarxiv-cs-ai
22 May 2026
Model Releases

Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Search

DGX agent

arXiv:2605.20244v1 Announce Type: cross Abstract: We present Lean Refactor, a plug-and-play retrieval-augmented agentic framework for multi-objective, controllable, and version-robust refactoring of L

model-releasesarxiv-cs-cl
21 May 2026
Model Releases

Learning Query-Aware Budget-Tier Routing for Runtime Agent Memory

DGX agent

arXiv:2602.06025v2 Announce Type: replace Abstract: Memory is increasingly central to Large Language Model (LLM) agents operating beyond a single context window, yet most existing systems rely on offl

model-releasesarxiv-cs-cl
21 May 2026
Model Releases

TRAM: Test-Time Risk Adaptation with Mixture of Agents

DGX agent

arXiv:2408.08812v2 Announce Type: replace Abstract: Deployed reinforcement learning agents often face safety requirements that are specified only after training, such as new hazard maps, revised risk

model-releasesarxiv-cs-lg
21 May 2026
Safety

Formal Skill: Programmable Runtime Skills for Efficient and Accurate LLM Agents

DGX agent

arXiv:2605.19604v1 Announce Type: new Abstract: Large Language Model (LLM) agents increasingly act inside real workspaces, where tools and skills determine whether model reasoning becomes reliable act

safetyarxiv-cs-ai
20 May 2026
Agents

KadiAssistant: A conversational AI Agent for information retrieval in Kadi4Mat

DGX agent

arXiv:2605.18850v1 Announce Type: cross Abstract: We introduce KadiAssistant, a privacy-by-design AI assistant integrated into the Kadi research data ecosystem, enabling researchers to efficiently acc

agentsarxiv-cs-ai
20 May 2026
Research

OpenComputer: Verifiable Software Worlds for Computer-Use Agents

DGX agent

arXiv:2605.19769v1 Announce Type: new Abstract: We present OpenComputer, a verifier-grounded framework for constructing verifiable software worlds for computer-use agents. OpenComputer integrates four

researcharxiv-cs-ai
20 May 2026
Model Releases

Physics-in-the-Loop: A Hybrid Agentic Architecture for Validated CAD Engineering Design

DGX agent

arXiv:2605.19717v1 Announce Type: new Abstract: Large Language Models (LLMs) can generate Computer-Aided Design (CAD), yet lack physical comprehension required for reliable engineering design. Instead

model-releasesarxiv-cs-cv
20 May 2026
Safety

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents

DGX agent

arXiv:2605.20061v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) is a promising paradigm for improving large language model (LLM) agents on long-horizon interactiv

safetyarxiv-cs-cl
20 May 2026
Agents

Robust Checkpoint Selection for Multimodal LLMs via Agentic Evaluation and Stability-Aware Ranking

DGX agent

arXiv:2605.18852v1 Announce Type: cross Abstract: Checkpoint selection for multimodal large language models (MLLMs) presents significant challenges when performance differentials are marginal and eval

agentsarxiv-cs-ai
20 May 2026
← Previous
1…9091929394…236
Next →