AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,153 results
Agents

Structuring versus Problematizing: How LLM-based Agents Scaffold Learning in Diagnostic Reasoning

DGX agent

arXiv:2604.09158v1 Announce Type: cross Abstract: Supporting students in developing diagnostic reasoning is a key challenge across educational domains. Novices often face cognitive biases such as prem

agentsarxiv-cs-ai
13 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Android Coach: Improve Online Agentic Training Efficiency with Single State Multiple Actions

DGX agent

arXiv:2604.07277v1 Announce Type: cross Abstract: Online reinforcement learning (RL) serves as an effective method for enhancing the capabilities of Android agents. However, guiding agents to learn th

safetyarxiv-cs-ai
10 Apr 2026
Model Releases

GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents

DGX agent

arXiv:2604.07429v1 Announce Type: new Abstract: Towards an embodied generalist for real-world interaction, Multimodal Large Language Model (MLLM) agents still suffer from challenging latency, sparse f

model-releasesarxiv-cs-cv
10 Apr 2026
Safety

Governed Capability Evolution for Embodied Agents: Safe Upgrade, Compatibility Checking, and Runtime Rollback for Embodied Capability Modules

DGX agent

arXiv:2604.08059v1 Announce Type: new Abstract: Embodied agents are increasingly expected to improve over time by updating their executable capabilities rather than rewriting the agent itself. Prior w

safetyarxiv-cs-ro
10 Apr 2026
Safety

Karma Mechanisms for Decentralised, Cooperative Multi Agent Path Finding

DGX agent

arXiv:2604.07970v1 Announce Type: cross Abstract: Multi-Agent Path Finding (MAPF) is a fundamental coordination problem in large-scale robotic and cyber-physical systems, where multiple agents must co

safetyarxiv-cs-ro
10 Apr 2026
Safety

KD-MARL: Resource-Aware Knowledge Distillation in Multi-Agent Reinforcement Learning

DGX agent

arXiv:2604.06691v1 Announce Type: new Abstract: Real world deployment of multi agent reinforcement learning MARL systems is fundamentally constrained by limited compute memory and inference time. Whil

safetyarxiv-cs-ai
10 Apr 2026
Agents

MAT-Cell: A Multi-Agent Tree-Structured Reasoning Framework for Batch-Level Single-Cell Annotation

DGX agent

arXiv:2604.06269v1 Announce Type: cross Abstract: Automated cellular reasoning faces a core dichotomy: supervised methods fall into the Reference Trap and fail to generalize to out-of-distribution cel

agentsarxiv-cs-ai
10 Apr 2026
Safety

Multi-agent Reach-avoid MDP via Potential Games and Low-rank Policy Structure

DGX agent

arXiv:2410.17690v2 Announce Type: replace-cross Abstract: We optimize finite horizon multi-agent reach-avoid Markov decision process (MDP) via local feedback policies. The global feedback polic

safetyarxiv-cs-ro
10 Apr 2026
Safety

TwinLoop: Simulation-in-the-Loop Digital Twins for Online Multi-Agent Reinforcement Learning

DGX agent

arXiv:2604.06610v1 Announce Type: cross Abstract: Decentralised online learning enables runtime adaptation in cyber-physical multi-agent systems, but when operating conditions change, learned policies

safetyarxiv-cs-ai
10 Apr 2026
Model Releases

VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics

DGX agent

arXiv:2604.06182v1 Announce Type: cross Abstract: Existing online benchmarks for mobile GUI agents remain largely app-centric and task-homogeneous, failing to reflect the diversity and instability of

model-releasesarxiv-cs-ai
10 Apr 2026
Agents

Escher-Loop: Mutual Evolution by Closed-Loop Self-Referential Optimization

DGX agent

arXiv:2604.23472v1 Announce Type: new Abstract: While recent autonomous agents demonstrate impressive capabilities, they predominantly rely on manually scripted workflows and handcrafted heuristics, i

agentsarxiv-cs-ai
28 Apr 2026
Agents

ChemWorld: Programmable Chemical Worlds for Controlled and Replayable Agent Experimentation

DGX agent

arXiv:2608.10792v1 Announce Type: new Abstract: Autonomous chemistry increasingly depends on environments in which agents can repeatedly act, observe, and adapt.Physical laboratories provide essential

agentsarxiv-cs-ai
12 Aug 2026
Model Releases

DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?

DGX agent

arXiv:2608.10366v1 Announce Type: new Abstract: Real-world data science involves long-horizon workflows that span data wrangling, exploration, modeling, visualization, and validation, and require coor

model-releasesarxiv-cs-ai
12 Aug 2026
Agents

Business Arena: Benchmarking LLM Agents in a Realistic Marketplace

DGX agent

arXiv:2608.08621v1 Announce Type: new Abstract: Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals, commit capital under uncertainty,

agentsarxiv-cs-ai
11 Aug 2026
Agents

FailForge: Distilling Procedural Competence from Persistent Failures into Code Agents

DGX agent

arXiv:2608.08570v1 Announce Type: new Abstract: Rejection sampling fine-tuning (RFT) is widely used to train code agents by generating trajectories on verifiable software engineering tasks, retaining

agentsarxiv-cs-ai
11 Aug 2026
Agents

Mendel Godel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution

DGX agent

arXiv:2608.07645v1 Announce Type: new Abstract: Self-improving coding agents that iteratively rewrite their own source code have demonstrated impressive performance on coding tasks. However, existing

agentsarxiv-cs-ai
11 Aug 2026
Model Releases

MetaSpace: Metamorphic Testing for Spatial Cognition in Embodied Agents

DGX agent

arXiv:2608.07533v1 Announce Type: new Abstract: An embodied agent is an intelligent entity that interacts with its environment through a physical body. Currently, the evaluation of embodied agents pri

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

MoRSE: Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts

DGX agent

arXiv:2608.09251v1 Announce Type: cross Abstract: Large language model-based multi-agent systems have recently shown strong potential for complex, long-horizon tasks. However, existing methods mainly

model-releasesarxiv-cs-ai
11 Aug 2026
Agents

Not an A11y: How Android Accessibility Exposes Mobile AI Agents to Indirect Prompt Injection

DGX agent

arXiv:2608.08939v1 Announce Type: new Abstract: The rise of autonomous AI agents represents a major paradigm shift in how users interact with mobile devices. Frameworks such as MobileRun and Mobile-Us

agentsarxiv-cs-ai
11 Aug 2026
Model Releases

SkillConsist: Detecting Inconsistencies in Agent Skills via Bidirectional Graph Alignment

DGX agent

arXiv:2608.07639v1 Announce Type: cross Abstract: Agent Skills provide reusable capabilities to LLM agents. Agent Skill inconsistencies can expose undisclosed dangerous behavior or cause wrong Skill s

model-releasesarxiv-cs-ai
11 Aug 2026
Agents

CAi Copilot: Reducing Operational Workload in Molecular Design through Intent-Driven Agentic Workflows

DGX agent

arXiv:2608.06961v1 Announce Type: new Abstract: Early-stage molecular design is an iterative process, not just a task of generating molecules. Researchers turn broad goals into design strategies, refi

agentsarxiv-cs-ai
10 Aug 2026
Model Releases

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing

DGX agent

arXiv:2608.07437v1 Announce Type: new Abstract: Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used to automate t

model-releasesarxiv-cs-ai
10 Aug 2026
Agents

Harnessing the Synergy between LLM Agents and Knowledge Graphs for Urban Socioeconomic Prediction

DGX agent

arXiv:2411.00028v3 Announce Type: replace-cross Abstract: Socioeconomic prediction aims to leverage various urban data to predict the socioeconomic indicators of regions such as population and commerc

agentsarxiv-cs-ai
10 Aug 2026
Agents

Online Security Learning in Cooperative Multi-Agent Systems under Hidden Byzantine Attacks

DGX agent

arXiv:2608.06520v1 Announce Type: new Abstract: We study online cooperative control of a multi-agent system under Byzantine attacks. Namely, an unknown, fixed subset of agents are Byzantine comprised

agentsarxiv-cs-lg
10 Aug 2026
Agents

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning

DGX agent

arXiv:2608.07371v1 Announce Type: cross Abstract: Recent agentic reinforcement learning methods use hindsight to complement sparse outcome rewards. However, a completed rollout can yield many such sig

agentsarxiv-cs-cl
10 Aug 2026
Agents

Agentic Nesting: A New Methodology for Existing Enterprise Application Integration and Services

DGX agent

arXiv:2608.05159v1 Announce Type: new Abstract: Enterprise operations extensively rely on multiple heterogeneous business systems and information applications, which also result in severe data silos a

agentsarxiv-cs-ai
7 Aug 2026
Agents

Beyond Relevance: Bayesian Evidence Acquisition for Agentic Whole-Slide Image Reasoning

DGX agent

arXiv:2608.05757v1 Announce Type: new Abstract: Whole-slide image (WSI) reasoning requires an agent to sequentially acquire visual evidence before answering a diagnostic question. Existing training-fr

agentsarxiv-cs-cv
7 Aug 2026
Agents

Hierarchical Server Architecture for Agentic Science

DGX agent

arXiv:2608.05332v1 Announce Type: cross Abstract: Agentic science is transforming the landscape of computational work, extending to scientific pipelines and workload managers. The workloads require sp

agentsarxiv-cs-ai
7 Aug 2026
Agents

Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents

DGX agent

arXiv:2608.06171v1 Announce Type: new Abstract: Web agents observe a browser through text, pixels, or both, and the choice is usually fixed once for all tasks. We measure six observation modes across

agentsarxiv-cs-cl
7 Aug 2026
Agents

The Vulnerability With No CVE: Managing Persistent Gaps Between Mandate and Authority in AI Coding Agents

DGX agent

arXiv:2608.05884v1 Announce Type: cross Abstract: Existing guidance identifies excessive agency, excessive permission, weak task-bound authorization, and inadequate agent controls as important risks.

agentsarxiv-cs-cl
7 Aug 2026
Agents

Communication-Enhanced Tutoring for Efficient Decentralized Multi-Agent Reinforcement Learning

DGX agent

arXiv:2508.13661v4 Announce Type: replace Abstract: Centralized Training with Decentralized Execution (CTDE) is the dominant paradigm in multi-agent reinforcement learning (MARL), enabling agents to a

agentsarxiv-cs-lg
6 Aug 2026
Agents

State2State: Environment-Derived Mid-Training for LLM Agents

DGX agent

arXiv:2608.04934v1 Announce Type: new Abstract: Training LLM agents commonly relies on supervised fine-tuning from expert trajectories or online reinforcement learning over human-specified tasks with

agentsarxiv-cs-cl
6 Aug 2026
Agents

stratum: A System Infrastructure for Massive Agent-Centric ML Workloads

DGX agent

arXiv:2603.03589v3 Announce Type: replace-cross Abstract: Recent advances in large language models (LLMs) transform how machine learning (ML) pipelines are developed and evaluated. LLMs enable a new t

agentsarxiv-cs-lg
6 Aug 2026
Agents

What Is a Skill Worth? Structure-Aware Shapley Valuation of Agent Skills

DGX agent

arXiv:2608.04562v1 Announce Type: new Abstract: Agent skills are increasingly optimized by automated feedback loops, producing long structured artifacts whose internal value remains unclear. We study

agentsarxiv-cs-ai
6 Aug 2026
Agents

CastFSR: A Fast--Slow--Reflect Agentic Reasoning Framework for Context-Aware Time Series Forecasting

DGX agent

arXiv:2608.03031v1 Announce Type: new Abstract: Time series forecasting is fundamental to decision-making in complex systems, where future dynamics are influenced not only by historical observations b

agentsarxiv-cs-ai
5 Aug 2026
Model Releases

PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents

DGX agent

arXiv:2608.04003v1 Announce Type: new Abstract: Recursive self-improvement requires agents to turn accumulated experience into better future behavior. Personal AI agents offer a concrete setting for s

model-releasesarxiv-cs-cl
5 Aug 2026
Agents

Traceable Multi-Agent System for Knowledge-Based Forecasting

DGX agent

arXiv:2608.03339v1 Announce Type: new Abstract: Enterprise forecasting increasingly relies on autonomous agents that interpret documents, search for data, generate code, and revise models. While this

agentsarxiv-cs-ai
5 Aug 2026
Agents

Verified Tool Calls Improve LLM Agent Reliability Under Non-Atomic Failures

DGX agent

arXiv:2608.02645v1 Announce Type: cross Abstract: Large Language Model (LLM) agents rely on external tools to perform multistage tasks. Existing agent frameworks typically assume that tool calls are a

agentsarxiv-cs-ai
5 Aug 2026
Model Releases

Control Under Compression: Reliability Frontiers for Tool-Using Agents

DGX agent

arXiv:2608.01056v1 Announce Type: cross Abstract: Tool-using language-model agents are governed not only by task prompts but also by persistent system-side instructions that specify tools, arguments,

model-releasesarxiv-cs-cl
4 Aug 2026
Safety

Cross-Domain Hybrid OPD for Generalizable Search Agents

DGX agent

arXiv:2608.02101v1 Announce Type: new Abstract: Recent advances in Reinforcement Learning (RL) have substantially improved the capabilities of autonomous search agents, enabling sophisticated planning

safetyarxiv-cs-cl
4 Aug 2026
Model Releases

CurveShift: Is Agent Progress Scalar? Separating Level from Shape

DGX agent

arXiv:2608.00355v1 Announce Type: new Abstract: Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

LoopsBench: From Harness Engineering to Loop Engineering in Benchmarking Coding Agent

DGX agent

arXiv:2608.00267v1 Announce Type: cross Abstract: Coding agent infrastructure is shifting from harness engineering toward loop engineering as coding agents are deployed for sustained long-horizon soft

model-releasesarxiv-cs-cl
4 Aug 2026
Safety

RMSWeb: Reflection, Failure-Mode Mining, and Salvage-DS for Web Agent Reinforcement Learning

DGX agent

arXiv:2608.00335v1 Announce Type: cross Abstract: Compact web agents can reduce deployment cost, but training them poses challenges in both data collection and post-SFT reinforcement learning (RL). Su

safetyarxiv-cs-cl
4 Aug 2026
Agents

SIEVE: Selective Integrity Verification and Escalation for Defending LLM Agents against Indirect Prompt Injection

DGX agent

arXiv:2512.06716v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly used as the core of agentic systems due to their strong reasoning, planning, and tool-use capabi

agentsarxiv-cs-cl
4 Aug 2026
Agents

Measuring Cognitive Engagement in Collaborative Discourse with an Extended ICAP Framework: Comparing Human Annotation, In-Context Learning, and Reflective LLM Agents

DGX agent

arXiv:2607.28651v1 Announce Type: cross Abstract: Collaboration supports learning and problem-solving, but its effectiveness depends on cognitive engagement during discourse. This study applies an ext

agentsarxiv-cs-cl
3 Aug 2026
Safety

Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember

DGX agent

arXiv:2607.29468v1 Announce Type: new Abstract: Self-play agents can generate training problems without questions from target benchmarks, but their curricula lack persistent state: failures affect gra

safetyarxiv-cs-ai
3 Aug 2026
Model Releases

IFCMemoryBench: Evaluating Long-Term Memory of LLM-Based Agents in BIM Information Retrieval

DGX agent

arXiv:2607.26072v1 Announce Type: cross Abstract: Long-term memory is becoming a core capability of LLM-based agents, but existing evaluations largely test conversational recall in open-domain or pers

model-releasesarxiv-cs-ai
31 Jul 2026
Model Releases

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents

DGX agent

arXiv:2607.26314v1 Announce Type: cross Abstract: Stealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelligence, is what separates sophisti

model-releasesarxiv-cs-ai
31 Jul 2026
← Previous
1…2425262728…233
Next →