AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,289 results
Tutorials

Do Agents Need to Plan Step-by-Step? Rethinking Planning Horizon in Data-Centric Tool Calling

DGX agent

arXiv:2605.08477v1 Announce Type: new Abstract: Explicit planning is a critical capability for LLM-based agents solving complex data-centric tasks, which require precise tool calling over external dat

tutorialsarxiv-cs-cl
12 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

Evolving-RL: End-to-End Optimization of Experience-Driven Self-Evolving Capability within Agents

DGX agent

arXiv:2605.10663v1 Announce Type: new Abstract: Experience-driven self-evolving agents aim to overcome the static nature of large language models by distilling reusable experience from past interactio

researcharxiv-cs-ai
12 May 2026
Local Ai

FlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage Orchestration

DGX agent

arXiv:2605.08520v1 Announce Type: new Abstract: LLM-based evolution has emerged as a promising way to improve agents by refining non-parametric artifacts, but its wall-clock cost remains a major bottl

local-aiarxiv-cs-lg
12 May 2026
Model Releases

Human-Inspired Memory Architecture for LLM Agents

DGX agent

arXiv:2605.08538v1 Announce Type: new Abstract: Current LLM agents lack principled mechanisms for managing persistent memory across long interaction horizons. We present a biologically-grounded memory

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

LLM Agents Already Know When to Call Tools -- Even Without Reasoning

DGX agent

arXiv:2605.09252v1 Announce Type: new Abstract: Tool-augmented LLM agents tend to call tools indiscriminately, even when the model can answer directly. Each unnecessary call wastes API fees and latenc

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

MemQ: Integrating Q-Learning into Self-Evolving Memory Agents over Provenance DAGs

DGX agent

arXiv:2605.08374v1 Announce Type: new Abstract: Episodic memory allows LLM agents to accumulate and retrieve experience, but current methods treat each memory independently, i.e., evaluating retrieval

model-releasesarxiv-cs-ai
12 May 2026
Safety

Position: Academic Conferences are Potentially Facing Denominator Gaming Caused by Fully Automated Scientific Agents

DGX agent

arXiv:2605.09915v1 Announce Type: cross Abstract: The implicit policy of maintaining relatively stable acceptance rates at top AI conferences, despite exponentially growing submissions, introduces a c

safetyarxiv-cs-ai
12 May 2026
Local Ai

PYTHALAB-MERA: Validation-Grounded Memory, Retrieval, and Acceptance Control for Frozen-LLM Coding Agents

DGX agent

arXiv:2605.08468v1 Announce Type: cross Abstract: Local LLM-based coding agents increasingly work in settings where correctness is earned through execution feedback, persistent state, and bounded repa

local-aiarxiv-cs-ai
12 May 2026
Safety

Skill-R1: Agent Skill Evolution via Reinforcement Learning

DGX agent

arXiv:2605.09359v1 Announce Type: cross Abstract: Agentic large language models often rely on skills, reusable natural language procedures that guide planning, action, and tool use. In practice, skill

safetyarxiv-cs-ai
12 May 2026
Model Releases

Why Retrying Fails: Context Contamination in LLM Agent Pipelines

DGX agent

arXiv:2605.08563v1 Announce Type: new Abstract: When an LLM agent fails a multi-step tool-augmented task and retries, the failed attempt typically remains in its context window -- contaminating the ne

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

End-to-end PDDL Planning with Hardcoded and Dynamic Agents

DGX agent

arXiv:2512.09629v2 Announce Type: replace Abstract: We present an end-to-end framework for planning supported by verifiers. An orchestrator receives a human specification written in natural language a

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

MAS-Algorithm: A Workflow for Solving Algorithmic Programming Problems with a Multi-Agent System

DGX agent

arXiv:2605.05949v2 Announce Type: replace Abstract: Algorithmic problem solving serves as a rigorous testbed for evaluating structured reasoning in AI coding systems, as it directly reflects a model's

model-releasesarxiv-cs-ai
11 May 2026
Safety

Same Signal, Opposite Meaning: Direction-Informed Adaptive Learning for LLM Agents

DGX agent

arXiv:2605.06908v1 Announce Type: cross Abstract: Adaptive test-time compute for LLM agents aims to invoke extra computation only when it improves performance. Existing methods typically use confidenc

safetyarxiv-cs-ai
11 May 2026
Safety

SHARP: A Self-Evolving Human-Auditable Rubric Policy for Financial Trading Agents

DGX agent

arXiv:2605.06822v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed for autonomous financial trading, a domain requiring continuous adaptation to noisy, non-stationa

safetyarxiv-cs-lg
11 May 2026
Applications

Weblica: Scalable and Reproducible Training Environments for Visual Web Agents

DGX agent

arXiv:2605.06761v1 Announce Type: new Abstract: The web is complex, open-ended, and constantly changing, making it challenging to scale training data for visual web agents. Existing data collection at

applicationsarxiv-cs-ai
11 May 2026
Safety

Scalable Multi Agent Diffusion Policies for Coverage Control

DGX agent

arXiv:2509.17244v2 Announce Type: replace Abstract: We propose MADP, a novel diffusion-model-based approach for collaboration in decentralized robot swarms. MADP leverages diffusion models to generate

safetyarxiv-cs-ro
7 May 2026
Model Releases

Storage Is Not Memory: A Retrieval-Centered Architecture for Agent Recall

DGX agent

arXiv:2605.04897v1 Announce Type: new Abstract: Extraction at ingestion is the wrong primitive for agent memory: content discarded before the query is known cannot be recovered at retrieval time. We p

model-releasesarxiv-cs-cl
7 May 2026
Safety

Agentic AI-Based Joint Computing and Networking via Mixture of Experts and Large Language Models

DGX agent

arXiv:2605.02911v1 Announce Type: new Abstract: Future sixth-generation (6G) mobile networks are envisioned to be equipped with a diverse set of powerful, yet highly specialized, optimization experts.

safetyarxiv-cs-lg
6 May 2026
Model Releases

DocSync: Agentic Documentation Maintenance via Critic-Guided Reflexion

DGX agent

arXiv:2605.02163v1 Announce Type: cross Abstract: Software documentation frequently drifts from executable logic as codebases evolve, creating technical debt that degrades maintainability and causes d

model-releasesarxiv-cs-ai
6 May 2026
Research

Empowering LLM Agents with Geospatial Awareness: Toward Grounded Reasoning for Wildfire Response

DGX agent

arXiv:2510.12061v2 Announce Type: replace Abstract: Effective disaster response is essential for safeguarding lives and property. Existing statistical approaches often lack semantic context, generaliz

researcharxiv-cs-ai
6 May 2026
Safety

Healthcare AI GYM for Medical Agents

DGX agent

arXiv:2605.02943v1 Announce Type: new Abstract: Clinical reasoning demands multi-step interactions -- gathering patient history, ordering tests, interpreting results, and making safe treatment decisio

safetyarxiv-cs-lg
6 May 2026
Model Releases

NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles

DGX agent

arXiv:2605.01847v1 Announce Type: new Abstract: Outcome-only evaluation under-specifies whether an evaluated agent profile preserves the commitments required to solve a multi-turn task coherently. Neu

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems

DGX agent

arXiv:2605.04018v1 Announce Type: new Abstract: Reasoning-intensive retrieval aims to surface evidence that supports downstream reasoning rather than merely matching topical similarity. This capabilit

model-releasesarxiv-cs-cl
6 May 2026
Safety

TRACE: A Metrologically-Grounded Engineering Framework for Trustworthy Agentic AI Systems in Operationally Critical Domains

DGX agent

arXiv:2605.03838v1 Announce Type: new Abstract: We introduce TRACE, a cross-domain engineering framework for trustworthy agentic AI in operationally critical domains. TRACE combines a four-layer refer

safetyarxiv-cs-cl
6 May 2026
Model Releases

ARIS: Agentic and Relationship Intelligence System for Social Robots

DGX agent

arXiv:2605.00943v1 Announce Type: new Abstract: Foundational models have advanced social robotics, enabling richer perception and communicative interaction with users. However, current systems still s

model-releasesarxiv-cs-ro
5 May 2026
Model Releases

CASE: An Agentic AI Framework for Enhancing Scam Intelligence in Digital Payments

DGX agent

arXiv:2508.19932v2 Announce Type: replace Abstract: The proliferation of digital payment platforms has transformed commerce, offering unmatched convenience and accessibility globally. However, this gr

model-releasesarxiv-cs-ai
5 May 2026
Safety

Compliance-Aware Agentic Payments on Stablecoin Rails

DGX agent

arXiv:2605.00071v1 Announce Type: cross Abstract: Agentic payment systems extend delegated action to financial transfers, but scaling them on stablecoin rails in regulated settings requires safeguards

safetyarxiv-cs-ai
5 May 2026
Model Releases

Coopetition-Gym v1: A Formally Grounded Platform for Mixed-Motive Multi-Agent Reinforcement Learning under Strategic Coopetition

DGX agent

arXiv:2605.02063v1 Announce Type: cross Abstract: We present Coopetition-Gym v1, a benchmark platform for mixed-motive multi-agent reinforcement learning under strategic coopetition. The platform comp

model-releasesarxiv-cs-lg
5 May 2026
Model Releases

FlexSQL: Flexible Exploration and Execution Make Better Text-to-SQL Agents

DGX agent

arXiv:2605.02815v1 Announce Type: new Abstract: Text-to-SQL over large analytical databases requires navigating complex schemas, resolving ambiguous queries, and grounding decisions in actual data. Mo

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

Gen-Searcher: Reinforcing Agentic Search for Image Generation

DGX agent

arXiv:2603.28767v2 Announce Type: replace Abstract: Recent image generation models have shown strong capabilities in generating high-fidelity and photorealistic images. However, they are fundamentally

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

PPO guided Agentic Pipeline for Adaptive Prompt Selection and Test Case Generation

DGX agent

arXiv:2605.00942v1 Announce Type: cross Abstract: Developing effective test cases capable of thoroughly exercising large-scale software systems is inherently difficult, especially if such systems have

model-releasesarxiv-cs-lg
5 May 2026
Safety

Rationality Measurement and Theory for Reinforcement Learning Agents

DGX agent

arXiv:2602.04737v2 Announce Type: replace Abstract: This paper proposes a suite of rationality measures and associated theory for reinforcement learning agents, a property increasingly critical yet ra

safetyarxiv-cs-lg
5 May 2026
Safety

Skills as Verifiable Artifacts: A Trust Schema and a Biconditional Correctness Criterion for Human-in-the-Loop Agent Runtimes

DGX agent

arXiv:2605.00424v1 Announce Type: cross Abstract: Agent skills -- structured packages of instructions, scripts, and references that augment a large language model (LLM) without modifying the model its

safetyarxiv-cs-ai
5 May 2026
Model Releases

A11y-Compressor: A Framework for Enhancing the Efficiency of GUI Agent Observations through Visual Context Reconstruction and Redundancy Reduction

DGX agent

arXiv:2605.00551v1 Announce Type: new Abstract: AI agents that interact with graphical user interfaces (GUIs) require effective observation representations for reliable grounding. The accessibility tr

model-releasesarxiv-cs-cl
4 May 2026
Model Releases

Learning from Supervision with Semantic and Episodic Memory: A Reflective Approach to Agent Adaptation

DGX agent

arXiv:2510.19897v2 Announce Type: replace Abstract: We investigate how agents built on pretrained large language models (LLMs) can learn target classification functions from labeled examples without p

model-releasesarxiv-cs-cl
4 May 2026
Safety

MemRouter: Memory-as-Embedding Routing for Long-Term Conversational Agents

DGX agent

arXiv:2605.00356v1 Announce Type: new Abstract: Long-term conversational agents must decide which turns to store in external memory, yet recent systems rely on autoregressive LLM generation at every t

safetyarxiv-cs-cl
4 May 2026
Safety

A High-Throughput Compute-Efficient POMDP Hide-And-Seek-Engine (HASE) for Multi-Agent Operations

DGX agent

arXiv:2604.27162v1 Announce Type: cross Abstract: Reinforcement Learning (RL) algorithms exhibit high sample complexity, particularly when applied to Decentralized Partially Observable Markov Decision

safetyarxiv-cs-lg
1 May 2026
Safety

Addressing the Reality Gap: A Three-Tension Framework for Agentic AI Adoption

DGX agent

arXiv:2604.27245v1 Announce Type: cross Abstract: Generative AI has rapidly entered education through free consumer tools, outpacing the ability of schools and universities to respond. Now a new wave

safetyarxiv-cs-ai
1 May 2026
Model Releases

MCPHunt: An Evaluation Framework for Cross-Boundary Data Propagation in Multi-Server MCP Agents

DGX agent

arXiv:2604.27819v1 Announce Type: new Abstract: Multi-server MCP agents create an information-flow control problem: faithful tool composition can turn individually benign read/write permissions into c

model-releasesarxiv-cs-ai
1 May 2026
Safety

Stable Behavior, Limited Variation: Persona Validity in LLM Agents for Urban Sentiment Perception

DGX agent

arXiv:2604.28048v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used as proxies for human perception in urban analysis, yet it remains unclear whether persona prompting p

safetyarxiv-cs-cl
1 May 2026
Model Releases

When Continual Learning Moves to Memory: A Study of Experience Reuse in LLM Agents

DGX agent

arXiv:2604.27003v1 Announce Type: cross Abstract: Memory-augmented LLM agents offer an appealing shortcut to continual learning: rather than updating model parameters, they accumulate experience in ex

model-releasesarxiv-cs-ai
1 May 2026
Safety

A Survey on the Safety and Security Threats of Computer-Using Agents: JARVIS or Ultron?

DGX agent

arXiv:2505.10924v4 Announce Type: replace-cross Abstract: Recently, AI-driven interactions with computing devices have advanced from basic prototype tools to sophisticated, LLM-based systems that emul

safetyarxiv-cs-ai
30 Apr 2026
Model Releases

ClawGym: A Scalable Framework for Building Effective Claw Agents

DGX agent

arXiv:2604.26904v1 Announce Type: cross Abstract: Claw-style environments support multi-step workflows over local files, tools, and persistent workspace states. However, scalable development around th

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

Deterministic Legal Agents: A Canonical Primitive API for Auditable Reasoning over Temporal Knowledge Graphs

DGX agent

arXiv:2510.06002v3 Announce Type: replace Abstract: In high-stakes legal domains, retrieval must preserve not only semantic relevance, but also the hierarchy, temporality, and causal provenance of leg

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

Why Search When You Can Transfer? Amortized Agentic Workflow Design from Structural Priors

DGX agent

arXiv:2604.25012v1 Announce Type: new Abstract: Automated agentic workflow design currently relies on per-task iterative search, which is computationally prohibitive and fails to reuse structural know

model-releasesarxiv-cs-lg
29 Apr 2026
Model Releases

Agentic Witnessing: Pragmatic and Scalable TEE-Enabled Privacy-Preserving Auditing

DGX agent

arXiv:2604.24203v1 Announce Type: cross Abstract: Auditing the semantic properties of proprietary data creates a fundamental tension: verification requires transparent access, while proprietary rights

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

GSAR: Typed Grounding for Hallucination Detection and Recovery in Multi-Agent LLMs

DGX agent

arXiv:2604.23366v1 Announce Type: new Abstract: Autonomous multi-agent LLM systems are increasingly deployed to investigate operational incidents and produce structured diagnostic reports. Their trust

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

LLM-Assisted Op-Amp Behavioral-Level Design via Agentic Human-Mimicking Reasoning

DGX agent

arXiv:2601.21321v2 Announce Type: replace Abstract: This paper proposes White-Op, an operational amplifier (op-amp) behavioral-level parameter design framework assisted by the human-mimicking reasonin

model-releasesarxiv-cs-ai
28 Apr 2026
← Previous
1…102103104105106…236
Next →