AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,289 results
Model Releases

CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?

DGX agent

arXiv:2605.16679v1 Announce Type: cross Abstract: End-to-end automation of realistic healthcare operations stresses three capabilities underrepresented in current benchmarks: policy density, decisions

model-releasesarxiv-cs-ai
19 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory

DGX agent

arXiv:2511.20857v2 Announce Type: replace-cross Abstract: Statefulness is essential for large language model (LLM) agents to perform long-term planning and problem-solving. This makes memory a critica

model-releasesarxiv-cs-ai
19 May 2026
Safety

Helping Customers in Distress: An LLM-powered Agent that Converses, Probes, and Routes

DGX agent

arXiv:2605.16268v1 Announce Type: cross Abstract: Banks receive millions of reports of fraud, scams, and disputed transactions every year, making it challenging to accurately direct customers to the a

safetyarxiv-cs-ai
19 May 2026
Model Releases

KISS - Knowledge Infrastructure for Scientific Simulation: A Scaffolding for Agentic Earth Science

DGX agent

arXiv:2605.17856v1 Announce Type: new Abstract: Process-based simulation models encode decades of scientific understanding across the Earth sciences, yet the communities most exposed to climate risk a

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

LongMINT: Evaluating Memory under Multi-Target Interference in Long-Horizon Agent Systems

DGX agent

arXiv:2605.18565v1 Announce Type: cross Abstract: Real-world agents operate over long and evolving horizons, where information is repeatedly updated and may interfere across memories, requiring accura

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

OSWorld-Human: Benchmarking the Efficiency of Computer-Use Agents

DGX agent

arXiv:2506.16042v2 Announce Type: replace Abstract: Generative AI is being leveraged to solve a variety of computer-use tasks involving desktop applications. State-of-the-art systems have focused sole

model-releasesarxiv-cs-ai
19 May 2026
Local Ai

PROTEA: Offline Evaluation and Iterative Refinement for Multi-Agent LLM Workflows

DGX agent

arXiv:2605.18032v1 Announce Type: cross Abstract: Multi-agent LLM workflows -- systems composed of multiple role-specific LLM calls -- often outperform single-prompt baselines, but they remain difficu

local-aiarxiv-cs-ai
19 May 2026
Local Ai

R2V Agent: Teaching SLMs When to Ask for Help

DGX agent

arXiv:2605.16604v1 Announce Type: new Abstract: Efficient agentic systems should incur expensive frontier-model costs only on decisions where a cheaper local model is likely to fail. Existing LLM casc

local-aiarxiv-cs-lg
19 May 2026
Model Releases

SE-GA: Memory-Augmented Self-Evolution for GUI Agents

DGX agent

arXiv:2605.16883v1 Announce Type: new Abstract: Autonomous Graphical User Interface (GUI) agents often struggle with multi-step tasks due to constrained context windows and static policies that fail t

model-releasesarxiv-cs-lg
19 May 2026
Agents

STRIDE: A Self-Reflective Agent Framework for Reliable Automatic Equation Discovery

DGX agent

arXiv:2605.17790v1 Announce Type: new Abstract: LLM-based equation discovery offers a promising route to recovering symbolic laws from data, but many systems still rely on generation-centered loops th

agentsarxiv-cs-ai
19 May 2026
Model Releases

Verify-Gated Completion as Admission Control in a Governed Multi-Agent Runtime: A Bounded Architecture Case Study

DGX agent

arXiv:2605.17998v1 Announce Type: cross Abstract: As multi-agent systems move from short interactions to tool-using workflows with specialized roles and persistent state, completion becomes a runtime-

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games

DGX agent

arXiv:2605.17637v1 Announce Type: new Abstract: Coding agents are increasingly used as application builders, yet many evaluations still focus on source code, repository-level tests, or intermediate tr

model-releasesarxiv-cs-ai
19 May 2026
Agents

Agent4POI: Agentic Context-Conditioned Affordance Reasoning for Multimodal Point-of-Interest Recommendation

DGX agent

arXiv:2605.15203v1 Announce Type: cross Abstract: We introduce Agent4POI, the first POI recommendation framework that generates context-conditioned multimodal representations at recommendation time, r

agentsarxiv-cs-ai
18 May 2026
Safety

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs

DGX agent

arXiv:2605.15565v1 Announce Type: cross Abstract: Reinforcement learning (RL) is increasingly used to improve the reasoning, coding, and tool-use capabilities of large language models, but agentic RL

safetyarxiv-cs-ai
18 May 2026
Agents

Autonomous Intelligent Agents for Natural-Language-Driven Web Execution with Integrated Security Assurance

DGX agent

arXiv:2605.15281v1 Announce Type: cross Abstract: Modern web test suites rot. A UI refactor breaks locators, a timing change causes race conditions, and within weeks developers abandon the suite entir

agentsarxiv-cs-ai
18 May 2026
Model Releases

FORGE: Self-Evolving Agent Memory With No Weight Updates via Population Broadcast

DGX agent

arXiv:2605.16233v1 Announce Type: new Abstract: Can LLM agents improve decision-making through self-generated memory without gradient updates? We propose FORGE (Failure-Optimized Reflective Graduation

model-releasesarxiv-cs-ai
18 May 2026
Model Releases

Rule2DRC: Benchmarking LLM Agents for DRC Script Synthesis with Execution-Guided Test Generation

DGX agent

arXiv:2605.15669v1 Announce Type: new Abstract: Manufacturable chip layouts must satisfy thousands of geometry-based design rules, and design rule checking (DRC) enforces them by running executable DR

model-releasesarxiv-cs-lg
18 May 2026
Model Releases

SkillSmith: Compiling Agent Skills into Boundary-Guided Runtime Interfaces

DGX agent

arXiv:2605.15215v1 Announce Type: new Abstract: Recently, skills have been widely adopted in large language model (LLM)-based agent systems across various domains. In existing frameworks, skills are t

model-releasesarxiv-cs-ai
18 May 2026
Local Ai

UniShield: An Adaptive Multi-Agent Framework for Unified Forgery Image Detection and Localization

DGX agent

arXiv:2510.03161v2 Announce Type: replace-cross Abstract: With the rapid advancements in image generation, synthetic images have become increasingly realistic, posing significant societal risks, such

local-aiarxiv-cs-ai
18 May 2026
Model Releases

A Tutorial on Cognitive Biases in Agentic AI-Driven 6G Autonomous Networks

DGX agent

arXiv:2510.19973v4 Announce Type: replace-cross Abstract: The path to higher network autonomy in 6G lies beyond the mere optimization of key performance indicators (KPIs), requiring systems that perce

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Agentic Recommender System with Hierarchical Belief-State Memory

DGX agent

arXiv:2605.14401v1 Announce Type: cross Abstract: Memory-augmented LLM agents have advanced personalized recommendation, yet existing approaches universally adopt flat memory representations that conf

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing

DGX agent

arXiv:2605.14084v1 Announce Type: cross Abstract: Code agents must both reason over long-horizon repository state and obey strict tool-use protocols. In paired Instruct/Thinking checkpoints, these cap

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations

DGX agent

arXiv:2605.14498v1 Announce Type: new Abstract: Large Language Model (LLM) agents increasingly serve as personal assistants and workplace collaborators, where their utility depends on memory systems t

model-releasesarxiv-cs-cl
15 May 2026
Model Releases

Holistic Evaluation and Failure Diagnosis of AI Agents

DGX agent

arXiv:2605.14865v1 Announce Type: new Abstract: AI agents execute complex multi-step processes, but current evaluation falls short: outcome metrics report success or failure without explaining why, an

model-releasesarxiv-cs-ai
15 May 2026
Agents

Hypergraph Enterprise Agentic Reasoner over Heterogeneous Business Systems

DGX agent

arXiv:2605.14259v1 Announce Type: new Abstract: Applying Large Language Models (LLMs) to heterogeneous enterprise systems is hindered by hallucinations and failures in multi-hop, n-ary reasoning. Exis

agentsarxiv-cs-ai
15 May 2026
Local Ai

LEMON: Learning Executable Multi-Agent Orchestration via Counterfactual Reinforcement Learning

DGX agent

arXiv:2605.14483v1 Announce Type: new Abstract: Large language models (LLMs) have become a strong foundation for multi-agent systems, but their effectiveness depends heavily on orchestration design. A

local-aiarxiv-cs-ai
15 May 2026
Safety

MAPLE: Latent Multi-Agent Play for End-to-End Autonomous Driving

DGX agent

arXiv:2605.14201v1 Announce Type: cross Abstract: Vision-language-action (VLA) models are effective as end-to-end motion planners, but can be brittle when evaluated in closed-loop settings due to bein

safetyarxiv-cs-cv
15 May 2026
Model Releases

pi-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows

DGX agent

arXiv:2605.14678v1 Announce Type: new Abstract: The rise of personal assistant agents, e.g., OpenClaw, highlights the growing potential of large language models to support users across everyday life a

model-releasesarxiv-cs-ai
15 May 2026
Safety

Self-Distilled Agentic Reinforcement Learning

DGX agent

arXiv:2605.15155v1 Announce Type: cross Abstract: Reinforcement learning (RL) has emerged as a central paradigm for post-training LLM agents, yet its trajectory-level reward signal provides only coars

safetyarxiv-cs-ai
15 May 2026
Research

Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents

DGX agent

arXiv:2605.12894v1 Announce Type: new Abstract: Large Language Model (LLM) agents are increasingly deployed in settings where they interact with a wide variety of people, including users who are uncle

researcharxiv-cs-ai
14 May 2026
Agents

MARLIN: Multi-Agent Game-Theoretic Reinforcement Learning for Sustainable LLM Inference in Cloud Datacenters

DGX agent

arXiv:2605.13496v1 Announce Type: cross Abstract: Large Language Models (LLMs) have become increasingly prevalent in cloud-based platforms, propelled by the introduction of AI-based consumer and enter

agentsarxiv-cs-lg
14 May 2026
Agents

Mechanism Plausibility in Generative Agent-Based Modeling

DGX agent

arXiv:2605.12824v1 Announce Type: cross Abstract: Large language models (LLMs) can generate high-level diverse phenomena without explicitly programmed rules. This capability has led to their adoption

agentsarxiv-cs-ai
14 May 2026
Agents

Multi-Dimensional Behavioral Evaluation of Agentic Stock Prediction Systems Using Large Language Model Judges with Closed-Loop Reinforcement Learning Feedback

DGX agent

arXiv:2605.05739v2 Announce Type: replace-cross Abstract: Forecast evaluation in finance has relied on aggregate accuracy metrics and predictive-accuracy tests built on point-forecast errors. These in

agentsarxiv-cs-ai
14 May 2026
Model Releases

Plan Before You Trade: Inference-Time Optimization for RL Trading Agents

DGX agent

arXiv:2605.12653v1 Announce Type: cross Abstract: Reinforcement learning agents for portfolio management are typically trained and deployed as static policies, with no mechanism for using price foreca

model-releasesarxiv-cs-ai
14 May 2026
Safety

ScioMind: Cognitively Grounded Multi-Agent Social Simulation with Anchoring-Based Belief Dynamics and Dynamic Profiles

DGX agent

arXiv:2605.13725v1 Announce Type: new Abstract: Large language model (LLM)-based multi-agent simulation offers a powerful testbed for studying social opinion dynamics. Yet current approaches often ado

safetyarxiv-cs-ai
14 May 2026
Safety

Think Twice, Act Once: Verifier-Guided Action Selection For Embodied Agents

DGX agent

arXiv:2605.12620v1 Announce Type: new Abstract: Building generalist embodied agents capable of solving complex real-world tasks remains a fundamental challenge in AI. Multimodal Large Language Models

safetyarxiv-cs-ai
14 May 2026
Model Releases

Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models

DGX agent

arXiv:2605.13119v1 Announce Type: cross Abstract: Vision-language-action (VLA) models are effective robot action executors, but they remain limited on long-horizon tasks due to the dual burden of exte

model-releasesarxiv-cs-ai
14 May 2026
Model Releases

Agent-BRACE: Decoupling Beliefs from Actions in Long-Horizon Tasks via Verbalized State Uncertainty

DGX agent

arXiv:2605.11436v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed on long-horizon tasks in partially observable environments, where they must act while inferring a

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

GRAFT-ATHENA: Self-Improving Agentic Teams for Autonomous Discovery and Evolutionary Numerical Algorithms

DGX agent

arXiv:2605.11117v1 Announce Type: new Abstract: Scientific discovery can be modeled as a sequence of probabilistic decisions that map physical problems to numerical solutions. Recent agentic AI system

model-releasesarxiv-cs-lg
13 May 2026
Safety

Predictive Maps of Multi-Agent Reasoning: A Successor-Representation Spectrum for LLM Communication Topologies

DGX agent

arXiv:2605.11453v1 Announce Type: cross Abstract: Practitioners deploying multi-agent large language model (LLM) systems must currently choose between communication topologies such as chain, star, mes

safetyarxiv-cs-lg
13 May 2026
Safety

Probabilistic Modeling of Latent Agentic Substructures in Deep Neural Networks

DGX agent

arXiv:2509.06701v2 Announce Type: replace Abstract: We develop a theory of intelligent agency grounded in probabilistic modeling for neural models. Agents are represented as outcome distributions with

safetyarxiv-cs-lg
13 May 2026
Agents

A Reflective Storytelling Agent for Older Adults: Integrating Argumentation Schemes and Argument Mining in LLM-Based Personalised Narratives

DGX agent

arXiv:2605.10531v1 Announce Type: new Abstract: This work investigates whether knowledge-driven large language model (LLM)-based storytelling can support purposeful narrative interaction with a digita

agentsarxiv-cs-ai
12 May 2026
Agents

A Versatile AI Agent for Rare Disease Diagnosis and Risk Gene Prioritization

DGX agent

arXiv:2605.06226v2 Announce Type: replace Abstract: Accurate and timely diagnosis is essential for effective treatment, particularly in the context of rare diseases. However, current diagnostic workfl

agentsarxiv-cs-ai
12 May 2026
Model Releases

Agentic Performance at the Edge: Insights from Benchmarking

DGX agent

arXiv:2605.10384v1 Announce Type: new Abstract: Agentic artificial intelligence (AI) is a natural fit for Internet of Things (IoT) and edge systems, but edge deployments are often constrained to model

model-releasesarxiv-cs-ai
12 May 2026
Agents

An agentic framework for gravitational-wave counterpart association in the multi-messenger era

DGX agent

arXiv:2605.10584v1 Announce Type: cross Abstract: With the detection of gravitational waves (GWs), multi-messenger astronomy has opened a new window for advancing our understanding of astrophysics, de

agentsarxiv-cs-ai
12 May 2026
Model Releases

Can Deep Research Agents Retrieve and Organize? Evaluating the Synthesis Gap with Expert Taxonomies

DGX agent

arXiv:2601.12369v3 Announce Type: replace Abstract: Deep Research Agents increasingly automate survey generation, yet whether they match human experts at retrieving essential papers and organizing the

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces

DGX agent

arXiv:2605.09497v1 Announce Type: new Abstract: Vision-language model (VLM) based web agents demonstrate impressive autonomous GUI interaction but remain vulnerable to deceptive interface elements. Ex

model-releasesarxiv-cs-ai
12 May 2026
Safety

Dynamic Skill Lifecycle Management for Agentic Reinforcement Learning

DGX agent

arXiv:2605.10923v1 Announce Type: cross Abstract: Large language model agents increasingly rely on external skills to solve complex tasks, where skills act as modular units that extend their capabilit

safetyarxiv-cs-cl
12 May 2026
← Previous
1…8182838485…236
Next →