AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,153 results
Model Releases

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents

DGX agent

arXiv:2605.23657v1 Announce Type: new Abstract: Skills, i.e., structured workflow instructions distilled for large language models (LLMs), are becoming an increasingly important mechanism for improvin

model-releasesarxiv-cs-cl
25 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

PACE: Two-Timescale Self-Evolution for Small Language Model Agents

DGX agent

arXiv:2605.23019v1 Announce Type: new Abstract: Deploying language-model agents in production often requires substantial compute and human effort to tune prompts, parsers, validators, and other compon

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

PhotoFlow: Agentic 3D Virtual Photography Missions

DGX agent

arXiv:2605.23771v1 Announce Type: cross Abstract: Virtual photography asks an agent to enter a prepared 3D scene with no preselected camera pose or reference image, infer a suitable shot from scene in

model-releasesarxiv-cs-ai
25 May 2026
Research

SciNet: Evaluating AI Agents in Relation-Aware Scientific Literature Retrieval

DGX agent

arXiv:2601.03260v2 Announce Type: replace-cross Abstract: AI agents have seen widespread adoption in information retrieval for scientific research, giving rise to tools such as Deep Research. However,

researcharxiv-cs-cl
25 May 2026
Model Releases

Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents

DGX agent

arXiv:2605.21768v1 Announce Type: new Abstract: Memory-augmented LLM agents enable interactions that extend beyond finite context windows by storing, updating, and reusing information across sessions.

model-releasesarxiv-cs-lg
23 May 2026
Model Releases

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI

DGX agent

arXiv:2603.14987v2 Announce Type: replace Abstract: Agentic AI systems increasingly act through tool-augmented, multi-step workflows whose failures (unsafe tool use, unauthorised actions, social harm)

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Declarative Data Services: Structured Agentic Discovery for Composing Data Systems

DGX agent

arXiv:2605.20690v1 Announce Type: new Abstract: Agentic discovery has shown that LLM-driven search can find novel algorithms, designs, and code under benchmark conditions. Translating the paradigm to

model-releasesarxiv-cs-ai
22 May 2026
Agents

MARS: Modular Agent with Reflective Search for Automated AI Research

DGX agent

arXiv:2602.02660v3 Announce Type: replace Abstract: A critical bottleneck in automating AI research is the execution of complex machine learning engineering (MLE) tasks. MLE differs from general softw

agentsarxiv-cs-ai
22 May 2026
Model Releases

APEX: Autonomous Policy Exploration for Self-Evolving LLM Agents

DGX agent

arXiv:2605.21240v1 Announce Type: new Abstract: LLM agents have shown strong performance across a wide range of complex tasks, including interactive environments that require long-horizon decision mak

model-releasesarxiv-cs-lg
21 May 2026
Model Releases

SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents

DGX agent

arXiv:2605.21384v1 Announce Type: cross Abstract: As long-horizon coding agents produce more code than any developer can review, oversight collapses onto a single surface: the automated test suite. Re

model-releasesarxiv-cs-cl
21 May 2026
Research

Weasel: Out-of-Domain Generalization for Web Agents via Importance-Diversity Data Selection

DGX agent

arXiv:2605.20291v1 Announce Type: new Abstract: Large language models (LLMs) have enabled web agents that follow natural language goals through multi-step browser interactions. However, agents fine-tu

researcharxiv-cs-lg
21 May 2026
Agents

What if Agents Could Imagine? Reinforcing Open-Vocabulary HOI Comprehension through Generation

DGX agent

arXiv:2602.11499v2 Announce Type: replace Abstract: Multimodal Large Language Models have shown promising capabilities in bridging visual and textual reasoning, yet their reasoning capabilities in Ope

agentsarxiv-cs-cv
21 May 2026
Model Releases

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema

DGX agent

arXiv:2605.21404v1 Announce Type: new Abstract: We read twelve well-known LLM agent benchmark papers and recorded, dimension by dimension, what each paper actually says about how its evaluation was ru

model-releasesarxiv-cs-lg
21 May 2026
Agents

A Multi-Agent Framework for Feature-Constrained Difficulty Control in Reading Comprehension Item Generation

DGX agent

arXiv:2605.19316v1 Announce Type: new Abstract: Recent studies in difficulty-controlled reading comprehension item generation have leveraged large language models (LLMs) to produce items by adjusting

agentsarxiv-cs-cl
20 May 2026
Model Releases

OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments

DGX agent

arXiv:2605.18758v1 Announce Type: cross Abstract: Current benchmarks for graphical user interface (GUI) agents predominantly rely on static screenshots. However, real-world smartphone interaction rout

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

RecoAtlas: From Semantic Plausibility to Set-Level Utility in LLM Recommendation Agents

DGX agent

arXiv:2605.18805v1 Announce Type: cross Abstract: LLM recommendation agents increasingly produce structured recommendation reports: sets of items accompanied by natural-language justifications. Yet ex

model-releasesarxiv-cs-ai
20 May 2026
Agents

Agentic Pipeline for Self-Synchronized Multiview Joint Angle Monitoring in Uncalibrated Environments

DGX agent

arXiv:2605.16419v1 Announce Type: cross Abstract: Kinematic monitoring plays a critical role in long-term rehabilitation for patients with spinal cord injury (SCI), where multi-view markerless motion

agentsarxiv-cs-ai
19 May 2026
Model Releases

ClawArena: Benchmarking AI Agents in Evolving Information Environments

DGX agent

arXiv:2604.04202v2 Announce Type: replace-cross Abstract: AI agents deployed as persistent assistants must maintain correct beliefs as their information environment evolves. In practice, evidence is s

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

ContraFix: Agentic Vulnerability Repair via Differential Runtime Evidence and Skill Reuse

DGX agent

arXiv:2605.17450v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly used for automated vulnerability repair (AVR), where repository-level reasoning enables them to ins

model-releasesarxiv-cs-ai
19 May 2026
Agents

DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents

DGX agent

arXiv:2605.17439v1 Announce Type: cross Abstract: Evaluating LLM-generated interactive software requires execution in addition to static analysis. The key difficulty is that correctness is a graph-lev

agentsarxiv-cs-ai
19 May 2026
Model Releases

Evaluating Cognitive Age Alignment in Interactive AI Agents

DGX agent

arXiv:2605.17894v1 Announce Type: new Abstract: While agentic AI and its core multimodal large language models (MLLMs) have demonstrated remarkable promise in language and visual reasoning across doma

model-releasesarxiv-cs-ai
19 May 2026
Safety

Natural-Language Agent Harnesses

DGX agent

arXiv:2603.25723v2 Announce Type: replace-cross Abstract: Agent performance is strongly shaped by the surrounding harness: the external execution system around a model that organizes a task run. Yet t

safetyarxiv-cs-ai
19 May 2026
Safety

Position: A Three-Layer Probabilistic Assume-Guarantee Architecture Is Structurally Required for Safe LLM Agent Deployment

DGX agent

arXiv:2605.18672v1 Announce Type: new Abstract: This position paper argues that enforcing LLM agent safety within a single abstraction layer is not merely suboptimal but categorically insufficient for

safetyarxiv-cs-ai
19 May 2026
Safety

PQR: A Framework to Generate Diverse and Realistic User Queries that Elicit QA Agent Failures

DGX agent

arXiv:2605.16551v1 Announce Type: new Abstract: Evaluating LLM-based agents remains challenging because identifying meaningful failure cases often requires substantial human effort to design realistic

safetyarxiv-cs-cl
19 May 2026
Safety

Reliability and Effectiveness of Autonomous AI Agents in Supply Chain Management

DGX agent

arXiv:2605.17036v1 Announce Type: new Abstract: This paper studies autonomous generative AI agents in multi-echelon supply chains using the MIT Beer Game. We identify four inference-time levers that s

safetyarxiv-cs-ai
19 May 2026
Model Releases

Responsible Agentic AI Requires Explicit Provenance

DGX agent

arXiv:2605.17169v1 Announce Type: new Abstract: Agentic AI is rapidly proliferating across diverse real-world domains such as software engineering, yet public trust has not kept pace. The central reas

model-releasesarxiv-cs-ai
19 May 2026
Research

Scalable Environments Drive Generalizable Agents

DGX agent

arXiv:2605.18181v1 Announce Type: new Abstract: Generalizable agents should adapt to diverse tasks and unseen environments beyond their training distribution. This position paper argues that such gene

researcharxiv-cs-ai
19 May 2026
Safety

TClone: Low-Latency Forking of Live GUI Environments for Computer-Use Agents

DGX agent

arXiv:2605.17320v1 Announce Type: cross Abstract: Computer-use agents increasingly operate inside live personal workspaces, where their actions can modify files, applications, GUI state, credentials,

safetyarxiv-cs-ai
19 May 2026
Safety

The Capability Paradox: How Smarter Auditors Make Multi-Agent Systems Less Secure

DGX agent

arXiv:2605.17480v1 Announce Type: new Abstract: Multi-agent systems extend large language models (LLMs) by decomposing tasks among specialized agents, but their distributed decision process creates ne

safetyarxiv-cs-ai
19 May 2026
Model Releases

TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents

DGX agent

arXiv:2605.16909v1 Announce Type: new Abstract: Tool-using agents are increasingly expected to operate across realistic professional workflows, where they must interpret multimodal inputs, coordinate

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback

DGX agent

arXiv:2605.17453v1 Announce Type: cross Abstract: Tool-using LLM agents increasingly rely on external tools to make consequential decisions, yet most existing agent-security benchmarks and defenses im

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

ALSO: Adversarial Online Strategy Optimization for Social Agents

DGX agent

arXiv:2605.15768v1 Announce Type: new Abstract: Social simulation provides a compelling testbed for studying social intelligence, where agents interact through multi-turn dialogues under evolving cont

model-releasesarxiv-cs-ai
18 May 2026
Agents

CodeDistiller: Automatically Generating Code Libraries for Scientific Coding Agents

DGX agent

arXiv:2512.01089v2 Announce Type: replace Abstract: Automated Scientific Discovery (ASD) systems can help automatically generate and run code-based experiments, but their capabilities are limited by t

agentsarxiv-cs-ai
18 May 2026
Agents

Detecting Privilege Escalation in Polyglot Microservices via Agentic Program Analysis

DGX agent

arXiv:2605.15569v1 Announce Type: cross Abstract: Microservices are widely adopted in modern cloud systems due to their scalability and fault tolerance. However, microservice architectures introduce s

agentsarxiv-cs-ai
18 May 2026
Model Releases

ShopGym: An Integrated Framework for Realistic Simulation and Scalable Benchmarking of E-Commerce Web Agents

DGX agent

arXiv:2605.16116v1 Announce Type: new Abstract: Developing and evaluating e-commerce web agents requires environments that preserve meaningful task structure while enabling controllable, reproducible,

model-releasesarxiv-cs-ai
18 May 2026
Model Releases

STAR: A Stage-attributed Triage and Repair framework for RCA Agents in Microservices

DGX agent

arXiv:2605.15581v1 Announce Type: new Abstract: LLM-based root cause analysis (RCA) agents have recently emerged as a promising paradigm for incident diagnosis in microservice AIOps. However, their re

model-releasesarxiv-cs-ai
18 May 2026
Model Releases

AgentTrap: Measuring Runtime Trust Failures in Third-Party Agent Skills

DGX agent

arXiv:2605.13940v1 Announce Type: cross Abstract: Third-party skills are becoming the package ecosystem for LLM agents. They package natural-language instructions, helper scripts, templates, documents

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Are Agents Ready to Teach? A Multi-Stage Benchmark for Real-World Teaching Workflows

DGX agent

arXiv:2605.14322v1 Announce Type: new Abstract: Language agents are increasingly deployed in complex professional workflows, with tutoring emerging as a particularly high-stakes capability that remain

model-releasesarxiv-cs-ai
15 May 2026
Agents

ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both

DGX agent

arXiv:2605.15198v1 Announce Type: cross Abstract: Visual reasoning, often interleaved with intermediate visual states, has emerged as a promising direction in the field. A straightforward approach is

agentsarxiv-cs-ai
15 May 2026
Agents

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents

DGX agent

arXiv:2510.02837v2 Announce Type: replace Abstract: Although recent tool-augmented benchmarks involve complex requests, evaluation remains limited to answer matching, neglecting critical trajectory as

agentsarxiv-cs-ai
15 May 2026
Agents

Contestable Multi-Agent Debate with Arena-based Argumentative Computation for Multimedia Verification

DGX agent

arXiv:2605.14495v1 Announce Type: cross Abstract: Multimedia verification requires not only accurate conclusions but also transparent and contestable reasoning. We propose a contestable multi-agent fr

agentsarxiv-cs-ai
15 May 2026
Agents

IFPV: An Integrated Multi-Agent Framework for Generative Operational Planning and High-Fidelity Plan Verification

DGX agent

arXiv:2605.14851v1 Announce Type: cross Abstract: Operational plan generation and verification are critical for modern complex and rapidly changing battlefield environments, yet traditional generation

agentsarxiv-cs-ai
15 May 2026
Agents

MediaClaw: Multimodal Intelligent-Agent Platform Technical Report

DGX agent

arXiv:2605.14771v1 Announce Type: new Abstract: MediaClaw is a multimodal agent platform built on the OpenClaw ecosystem. Its core design follows a three-layer architecture of unified abstraction, plu

agentsarxiv-cs-ai
15 May 2026
Model Releases

Sheaf-Theoretic Transport and Obstruction for Detecting Scientific Theory Shift in AI Agents

DGX agent

arXiv:2605.14033v1 Announce Type: new Abstract: Scientific theory shift in AI agents requires more than fitting equations to data. An artificial scientific agent must detect whether an existing repres

model-releasesarxiv-cs-ai
15 May 2026
Safety

Temporal Fair Division in Multi-Agent Systems: From Precise Alternation Metrics to Scalable Coordination Proxies

DGX agent

arXiv:2605.14879v1 Announce Type: cross Abstract: A plethora real-world environments require agents to compete repeatedly for the same limited resource, calling for a temporal notion of fairness judge

safetyarxiv-cs-lg
15 May 2026
Agents

A Multi-Agent Orchestration Framework for Venture Capital Due Diligence

DGX agent

arXiv:2605.13110v1 Announce Type: cross Abstract: We present a fully automated multi-agent framework for corporate due diligence and market analysis in venture capital. The system runs on an event-dri

agentsarxiv-cs-ai
14 May 2026
Agents

AI co-mathematician: Accelerating mathematicians with agentic AI

DGX agent

arXiv:2605.06651v2 Announce Type: replace Abstract: We introduce the AI co-mathematician, a workbench for mathematicians to interactively leverage AI agents to pursue open-ended research. The AI co-ma

agentsarxiv-cs-ai
14 May 2026
Local Ai

CANTANTE: Optimizing Agentic Systems via Contrastive Credit Attribution

DGX agent

arXiv:2605.13295v1 Announce Type: cross Abstract: LLM-based multi-agent systems have demonstrated strong performance across complex real-world tasks, such as software engineering, predictive modeling,

local-aiarxiv-cs-ai
14 May 2026
← Previous
1…5960616263…233
Next →