AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,289 results
Model Releases

Skill Availability and Presentation Granularity in Large-Language-Model Agents: A Controlled SkillsBench Study

DGX agent

arXiv:2605.31408v1 Announce Type: cross Abstract: Skill documents provide procedural knowledge to large-language-model agents at inference time. This article studies whether the presentation granulari

model-releasesarxiv-cs-ai
1 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling

DGX agent

arXiv:2605.29697v1 Announce Type: new Abstract: In Agentic Search, trajectory-level outcome rewards fail to quantify the behavioral contributions of individual steps, while existing step-level reward

safetyarxiv-cs-ai
29 May 2026
Safety

GAPD: Gold-Action Policy Distillation for Agentic Reinforcement Learning in Knowledge Base Question Answering

DGX agent

arXiv:2605.29584v1 Announce Type: new Abstract: Reinforcement learning (RL) is a natural fit for agentic knowledge base question answering (KBQA), where a model must issue executable actions, observe

safetyarxiv-cs-cl
29 May 2026
Model Releases

GroundAct: Can LLM Agents Ground Actions in Environmental States?

DGX agent

arXiv:2508.05614v2 Announce Type: replace-cross Abstract: LLM agents achieve 85-96% success on tasks where instructions fully specify the action, but drop to 29-53% when action feasibility depends on

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

InsightEval: An Expert-Curated Benchmark for Assessing Insight Discovery in LLM-Driven Data Agents

DGX agent

arXiv:2511.22884v2 Announce Type: replace Abstract: Data analysis has become an indispensable part of scientific research. To discover the latent knowledge and insights hidden within massive datasets,

model-releasesarxiv-cs-ai
29 May 2026
Local Ai

Locally Coherent, Globally Incoherent: Bounding Compositional Incoherence in Multi-Component LLM Agents

DGX agent

arXiv:2605.30335v1 Announce Type: new Abstract: Multi-component LLM agents assemble probabilistic claims from components that each see only part of a joint problem; the composition can violate basic p

local-aiarxiv-cs-ai
29 May 2026
Safety

Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents

DGX agent

arXiv:2605.29447v1 Announce Type: cross Abstract: While GUI agents have advanced rapidly, they often lack the robustness to recover from their own errors, hindering real-world deployment. To bridge th

safetyarxiv-cs-cl
29 May 2026
Model Releases

Scaling Laws for Agent Harnesses via Effective Feedback Compute

DGX agent

arXiv:2605.29682v1 Announce Type: new Abstract: Agent harnesses increasingly determine the performance of language-model systems by deciding how models call tools, receive feedback, verify intermediat

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

SURGENT: A Surgical Multi-Agent Assistance System Across the Perioperative Workflow

DGX agent

arXiv:2605.29368v1 Announce Type: cross Abstract: The intricate nature of modern surgical care necessitates intelligent systems that can synthesize extensive patient records, support collaborative dec

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Beyond One Path: Evaluating and Enhancing Divergent Thinking in Interactive LLM Agents

DGX agent

arXiv:2605.28465v1 Announce Type: new Abstract: Divergent thinking is a core dimension of creativity, yet existing evaluations of Large Language Models (LLMs) treat them as single-turn text generation

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

ClinicalAgents: Multi-Agent Orchestration for Clinical Decision Making with Dual-Memory

DGX agent

arXiv:2603.26182v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have demonstrated potential in healthcare, they often struggle with the complex, non-linear reasoning required fo

model-releasesarxiv-cs-cl
28 May 2026
Safety

CPPO: Contrastive Perception Policy Optimization for VLM Agents

DGX agent

arXiv:2601.00501v2 Announce Type: replace Abstract: We introduce CPPO, a Contrastive Perception Policy Optimization method for finetuning vision--language models (VLMs). Reliable perception is a core

safetyarxiv-cs-cv
28 May 2026
Safety

Diagnosing Live Within-Policy Instruction Conflicts in LLM Agents with Witnessed Resolution Profiles

DGX agent

arXiv:2605.27784v1 Announce Type: new Abstract: LLM agents are governed by long-lived natural-language prompt policies, but individually reasonable standing rules can interact in uninspected ways. We

safetyarxiv-cs-ai
28 May 2026
Safety

Evaluating the Realism of LLM-powered Social Agents: A Case Study of Reactions to Spanish Online News

DGX agent

arXiv:2605.28598v1 Announce Type: cross Abstract: LLM-powered social agents are increasingly used to simulate online social behavior, yet their realism remains difficult to validate. Existing work has

safetyarxiv-cs-ai
28 May 2026
Model Releases

Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills

DGX agent

arXiv:2604.05333v3 Announce Type: replace Abstract: Modern LLM agents increasingly rely on reusable skills, and as they interact with personal applications, web browsers, and other interfaces, skill l

model-releasesarxiv-cs-ai
28 May 2026
Safety

ICAN-Deploy: Identity-Stable Canary Deployment for Safety-Critical Embodied Agents

DGX agent

arXiv:2605.28097v1 Announce Type: new Abstract: Canary deployment routes a fraction of traffic to a new software version, monitors metrics, and rolls back on regression. Mainstream controllers (Argo R

safetyarxiv-cs-ro
28 May 2026
Safety

Identifying and Mitigating Bottlenecks in Role-Playing Agents: A Systematic Study of Disentangling Character Profile Axes

DGX agent

arXiv:2601.04716v3 Announce Type: replace Abstract: While Large Language Model (LLM) role-playing agents have advanced rapidly, it remains unclear which profile elements genuinely drive role-playing q

safetyarxiv-cs-cl
28 May 2026
Model Releases

LCO: LLM-based Constraint Optimization for Safer Agentic LLMs in Real-world Tasks

DGX agent

arXiv:2605.27375v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly acting as autonomous agents, but their continuous interaction with the environment can lead to in-context

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning

DGX agent

arXiv:2605.27960v1 Announce Type: new Abstract: Despite their popularity and success, Multimodal Large Language Models (MLLMs) often struggle to interpret images accurately, which limits their reasoni

model-releasesarxiv-cs-cv
28 May 2026
Safety

SKILLC: Learning Autonomous Skill Internalization in LLM Agents via Contrastive Credit Assignment

DGX agent

arXiv:2605.27899v1 Announce Type: new Abstract: Structured skill prompts improve exploration in long-horizon agentic reinforcement learning (RL). Skill-augmented RL methods retain external skills at i

safetyarxiv-cs-ai
28 May 2026
Model Releases

AGORA: Adapter-Grounded Observation-Action Retention for Inference-Free Prompt Compression in LLM Agents

DGX agent

arXiv:2605.26596v1 Announce Type: new Abstract: The token-level extractive compressors widely used for general LM context are structurally inappropriate for LLM agents: across 17 (env, backbone, metho

model-releasesarxiv-cs-ai
27 May 2026
Safety

AlignEvoSkill: Towards Knowledge-Aware and Task-Aligned Agent Skill Evolution

DGX agent

arXiv:2506.23149v2 Announce Type: replace Abstract: Reusable skills play a key role in improving LLM-based agents, but existing skill-evolution methods often fail to ensure that evolved skills both co

safetyarxiv-cs-cl
27 May 2026
Safety

Foundations of a Time-Consistent Counterfactual Actuarial Runtime for Autonomous AI Agents

DGX agent

arXiv:2605.26508v1 Announce Type: cross Abstract: We propose a foundational runtime actuarial layer for autonomous AI agents in which every side-effect-bearing action carries a time-consistent, counte

safetyarxiv-cs-ai
27 May 2026
Safety

From Norms to Indicators (N2I-RAG): An Agentic Retrieval-Augmented Generation Framework for Legal Indicator Computation

DGX agent

arXiv:2605.26926v1 Announce Type: new Abstract: Computing legal indicators from normative texts is a key task in legal monitoring and policy evaluation, but presents significant challenges due to the

safetyarxiv-cs-ai
27 May 2026
Model Releases

It's Not the Capability: Harness Sensitivity Is Non-Monotone Across LLM Agent Tiers

DGX agent

arXiv:2605.26731v1 Announce Type: new Abstract: A prevalent assumption in LLM agent deployment holds that more structured harnesses universally improve reliability, and that higher-capability models n

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Multi-Agent Causal Discovery Using Large Language Models

DGX agent

arXiv:2407.15073v4 Announce Type: replace Abstract: Causal discovery aims to identify causal relationships between variables and is a fundamental problem across the sciences. Traditional statistical c

model-releasesarxiv-cs-ai
27 May 2026
Research

Natural Language Query to Configuration for Retrieval Agents

DGX agent

arXiv:2605.27361v1 Announce Type: new Abstract: Modern retrieval agents expose many configuration choices -- LLM, retriever, number of documents, number of hops, and synthesis strategy -- each shaping

researcharxiv-cs-ai
27 May 2026
Hardware

Towards Feedback-to-Plan Decisions for Self-Evolving LLM Agents in CUDA Kernel Generation

DGX agent

arXiv:2605.26720v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong empirical gains as self-evolving agents for CUDA kernel generation, driven by feedback-conditioned planni

hardwarearxiv-cs-ai
27 May 2026
Model Releases

Automated Benchmark Auditing for AI Agents and Large Language Models

DGX agent

arXiv:2605.26079v1 Announce Type: new Abstract: Modern AI benchmarks operate at a complexity that outpaces traditional verification methods. Tasks authored by domain experts often contain implicit ass

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Code2UML: Agentic LLMs with context engineering for scalable software visualization

DGX agent

arXiv:2605.24453v1 Announce Type: cross Abstract: Large Language Model (LLM)-based code analysis tools are adopted to automate software documentation tasks. However, the scalability of these approache

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions

DGX agent

arXiv:2605.24279v1 Announce Type: new Abstract: A frontier language model's acknowledged 'helpful programming assistant' persona does not survive long agentic-coding sessions in the deployment regime

model-releasesarxiv-cs-cl
26 May 2026
Safety

Contractual Skills: A GovernSpec Design Framework for Enterprise AI Agents

DGX agent

arXiv:2605.22634v2 Announce Type: replace-cross Abstract: Skills have become a practical packaging mechanism for agent instructions, workflows, scripts, and reference materials. In enterprise settings

safetyarxiv-cs-ai
26 May 2026
Model Releases

From Facts to Insights: A Persona-Driven Dual Memory Framework and Dataset for Role-Playing Agents

DGX agent

arXiv:2605.25693v1 Announce Type: new Abstract: While role-playing agents excel in short-term interactions, long-term conversations overwhelm context windows, motivating external memory frameworks. Cu

model-releasesarxiv-cs-cl
26 May 2026
Local Ai

LLM Agent Based Renewable Energy Forecasting Using Edge and IoT Data A Review of Solar Wind Weather and Grid Aware Decision Support

DGX agent

arXiv:2605.25141v1 Announce Type: cross Abstract: Reliable forecasting of renewable energy generation is a foundational requirement for grid stability energy trading battery scheduling and carbon awar

local-aiarxiv-cs-ai
26 May 2026
Safety

Micro-Swarm Locomotion Optimization in Dynamic Flow using Multi-Objective Multi-Agent Reinforcement Learning

DGX agent

arXiv:2605.25025v1 Announce Type: new Abstract: Coordinating micro-robotic swarms in physiologically realistic, time-dependent fluid environments remains an unsolved challenge for biomedical and envir

safetyarxiv-cs-ro
26 May 2026
Model Releases

Neural Router: Semantic Content Matching for Agentic AI

DGX agent

arXiv:2605.25701v1 Announce Type: cross Abstract: Large language models (LLMs) can serve as the semantic-matching engine of a content-based publish/subscribe broker for agentic AI across the edge-clou

model-releasesarxiv-cs-cl
26 May 2026
Safety

Operationalizing Reconstructive Authority: Runtime Construction, Dependency Resolution, and Execution Gating in Autonomous Agent Systems

DGX agent

arXiv:2605.23935v1 Announce Type: new Abstract: Autonomous agent systems fail not only due to incorrect decisions, but due to executing decisions whose authority no longer holds at runtime. Prior work

safetyarxiv-cs-ai
26 May 2026
Safety

ProActor: Timing-Aware Reinforcement Learning for Proactive Task Scheduling Agents

DGX agent

arXiv:2605.24900v1 Announce Type: new Abstract: Proactive task-oriented agents must autonomously anticipate user needs, identify actionable opportunities, and trigger software actions at appropriate m

safetyarxiv-cs-ai
26 May 2026
Local Ai

DART: Semantic Recoverability for Structured Tool Agents

DGX agent

arXiv:2605.23311v1 Announce Type: new Abstract: When a structured tool agent fails mid-execution, the runtime faces a dilemma: replaying the entire task is safe but wasteful, while restoring from a lo

local-aiarxiv-cs-ai
25 May 2026
Local Ai

PathNavigate: A Training-Free Pathology Agent with Surprise-Guided Scan and Shared Slide Memory for Whole-Slide Image VQA

DGX agent

arXiv:2605.23559v1 Announce Type: cross Abstract: Whole-slide image visual question answering (WSI-VQA) frames pathology as an extreme-context search problem: to answer a free-form clinical query, a s

local-aiarxiv-cs-ai
25 May 2026
Research

Dynamic Mixture of Latent Memories for Self-Evolving Agents

DGX agent

arXiv:2605.21951v1 Announce Type: new Abstract: Achieving self-evolution in intelligent agents requires the continual accumulation of new knowledge across changing task sequences without forgetting pr

researcharxiv-cs-lg
23 May 2026
Model Releases

MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents

DGX agent

arXiv:2602.13372v2 Announce Type: replace-cross Abstract: Evaluating moral alignment in agents navigating conflicting, hierarchically structured human norms is a critical challenge at the intersection

model-releasesarxiv-cs-lg
23 May 2026
Model Releases

Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines

DGX agent

arXiv:2605.20630v1 Announce Type: new Abstract: Industrial asset operations workflows are latency-sensitive because a single user query may require coordination over sensor data, work orders, failure

model-releasesarxiv-cs-ai
22 May 2026
Tutorials

Personality Engineering with AI Agents: A New Methodology for Negotiation Research

DGX agent

arXiv:2605.20554v1 Announce Type: new Abstract: According to canonical negotiation theory, people's success in a negotiation depends on how well they balance competing demands--empathizing and asserti

tutorialsarxiv-cs-ai
22 May 2026
Model Releases

RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems

DGX agent

arXiv:2510.13910v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) mitigates key limitations of Large Language Models (LLMs)-such as factual errors, outdated knowledge, and hallu

model-releasesarxiv-cs-cl
22 May 2026
Safety

Retaining Suboptimal Actions to Follow Shifting Optima in Multi-Agent Reinforcement Learning

DGX agent

arXiv:2602.17062v2 Announce Type: replace Abstract: Value decomposition is a core approach for cooperative multi-agent reinforcement learning (MARL). However, existing methods still rely on a single o

safetyarxiv-cs-ai
22 May 2026
Safety

Multi-Agent Reinforcement Learning for Safe Autonomous Driving Under Pedestrian Behavioral Uncertainty

DGX agent

arXiv:2605.20255v1 Announce Type: new Abstract: Simulation-based testing of self-driving cars (SDCs) typically relies on scripted or simplified pedestrian models that do not capture the heterogeneity

safetyarxiv-cs-lg
21 May 2026
Safety

STEAM: A Training-Free Congestion-Aware Enhancement Framework for Decentralized Multi-Agent Path Finding

DGX agent

arXiv:2605.20929v1 Announce Type: new Abstract: We propose STEAM (Spatial, Temporal, and Emergent congestion Awareness for MAPF), a training-free test-time enhancement framework for learning-based dec

safetyarxiv-cs-ro
21 May 2026
← Previous
1…100101102103104…236
Next →