AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,289 results
Model Releases

Measuring Safety Alignment Effects in Autonomous Security Agents

DGX agent

arXiv:2605.19722v1 Announce Type: cross Abstract: Do stock safety-aligned language models and their uncensored or abliterated derivatives behave differently when run as autonomous security agents? Sin

model-releasesarxiv-cs-ai
20 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Progressive Autonomy as Preference Learning: A Formalization of Trust Calibration for Agentic Tool Use

DGX agent

arXiv:2605.19151v1 Announce Type: new Abstract: We formalize trust calibration for agentic tool use (deciding when an automated agent's proposed action may execute autonomously versus require human ap

safetyarxiv-cs-ai
20 May 2026
Model Releases

Sequential Consensus for Multi-Agent LLM Debates: A Wald-SPRT compute governor with calibration-based failure detection

DGX agent

arXiv:2605.19193v1 Announce Type: new Abstract: Multi-agent LLM debate improves factuality and reasoning, but most recipes pick a fixed round count, over-spending on easy items and under-spending on h

model-releasesarxiv-cs-lg
20 May 2026
Model Releases

The World Won't Stay Still: Programmable Evolution for Agent Benchmarks

DGX agent

arXiv:2603.05910v2 Announce Type: replace Abstract: LLM-powered tool-calling agents fulfill user requests by interacting with environments, querying data, and invoking tools in a multi-turn process. Y

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Adversarial Agent Collaboration for Correctness Improvements of C to Safe Rust Translation

DGX agent

arXiv:2510.03879v3 Announce Type: replace-cross Abstract: Translating C to memory-safe languages, like Rust, prevents critical memory safety vulnerabilities that are prevalent in legacy C software. Ev

model-releasesarxiv-cs-ai
19 May 2026
Safety

AI Agents May Always Fall for Prompt Injections

DGX agent

arXiv:2605.17634v1 Announce Type: cross Abstract: Prompt injection is the most critical vulnerability in deployed AI agents. Despite recent progress, we show that the prevailing defense paradigm (data

safetyarxiv-cs-cl
19 May 2026
Local Ai

Aurora: Unified Video Editing with a Tool-Using Agent

DGX agent

arXiv:2605.18748v1 Announce Type: new Abstract: Recent video editing models have converged on a unified conditioning design: a single diffusion transformer jointly consumes text, source video, and ref

local-aiarxiv-cs-cv
19 May 2026
Model Releases

CVE-Factory: Scaling Expert-Level Agentic Tasks for Code Security Vulnerability

DGX agent

arXiv:2602.03012v2 Announce Type: replace-cross Abstract: Evaluating and improving the security capabilities of code agents requires high-quality, executable vulnerability tasks. However, existing wor

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

EvoMemBench: Benchmarking Agent Memory from a Self-Evolving Perspective

DGX agent

arXiv:2605.18421v1 Announce Type: cross Abstract: Recent benchmarks for Large Language Model (LLM) agents mainly evaluate reasoning, planning, and execution. However, memory is also essential for agen

model-releasesarxiv-cs-ai
19 May 2026
Agents

Learning from Self-Debate: Preparing Reasoning Models for Multi-Agent Debate

DGX agent

arXiv:2601.22297v2 Announce Type: replace Abstract: The reasoning abilities of large language models (LLMs) have been substantially improved by reinforcement learning with verifiable rewards (RLVR). A

agentsarxiv-cs-cl
19 May 2026
Agents

MA^{2}P: A Meta-Cognitive Autonomous Intelligent Agents Framework for Complex Persuasion

DGX agent

arXiv:2605.18572v1 Announce Type: new Abstract: Persuasive dialogue generation plays a vital role in decision-making, negotiation, counseling, and behavior change, yet it remains a challenging problem

agentsarxiv-cs-cl
19 May 2026
Agents

RAGA: Reading-And-Graph-building-Agent for Autonomous Knowledge Graph Construction and Retrieval-Augmented Generation

DGX agent

arXiv:2605.17072v1 Announce Type: new Abstract: Existing LLM-driven knowledge graph (KG) construction methods predominantly employ stateless batch processing pipelines, exhibiting structural deficienc

agentsarxiv-cs-ai
19 May 2026
Model Releases

SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors

DGX agent

arXiv:2605.16626v1 Announce Type: cross Abstract: Since autonomous coding agents generate complex behaviors at high-volume, we may want to use other LLMs to monitor actions to reduce the risk from dan

model-releasesarxiv-cs-ai
19 May 2026
Safety

State Contamination in Memory-Augmented LLM Agents

DGX agent

arXiv:2605.16746v1 Announce Type: new Abstract: LLM agents increasingly rely on persistent state, including transcripts, summaries, retrieved context, and memory buffers, to support long-horizon inter

safetyarxiv-cs-ai
19 May 2026
Model Releases

Supervising the search process produces reliable and generalizable information-seeking agents

DGX agent

arXiv:2502.13957v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are transforming web search by shifting from document ranking to synthesizing answers, and are increasingly deplo

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents

DGX agent

arXiv:2605.16282v1 Announce Type: cross Abstract: The rapid deployment of LLM-based autonomous agents has introduced safety risks that extend far beyond traditional LLM concerns, prompting a prolifera

model-releasesarxiv-cs-ai
19 May 2026
Agents

The End of Trust: How Agentic AI Breaks Security Assumptions

DGX agent

arXiv:2605.16436v1 Announce Type: cross Abstract: For decades, the security of digital interaction has rested on an unacknowledged economic constraint. Attackers faced a tradeoff between the fidelity

agentsarxiv-cs-ai
19 May 2026
Model Releases

When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents

DGX agent

arXiv:2601.17887v2 Announce Type: replace Abstract: Long-term memory enables large language model (LLM) agents to support personalized and sustained interactions. However, most work on personalized ag

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

A3D: Agentic AI flow for autonomous Accelerator Design

DGX agent

arXiv:2605.15237v1 Announce Type: cross Abstract: Accelerating applications through the design of hardware accelerators can significantly enhance system performance and energy efficiency. Despite adva

model-releasesarxiv-cs-ai
18 May 2026
Model Releases

Argus: Evidence Assembly for Scalable Deep Research Agents

DGX agent

arXiv:2605.16217v1 Announce Type: cross Abstract: Deep research agents have achieved remarkable progress on complex information seeking tasks. Even long ReAct style rollouts explore only a single traj

model-releasesarxiv-cs-ai
18 May 2026
Model Releases

CryptoBench: A Dynamic Benchmark for Expert-Level Evaluation of LLM Agents in Cryptocurrency

DGX agent

arXiv:2512.00417v5 Announce Type: replace Abstract: This paper introduces CryptoBench, the first expert-curated, dynamic benchmark designed to rigorously evaluate the real-world capabilities of Large

model-releasesarxiv-cs-cl
18 May 2026
Model Releases

FormulaCode: Evaluating Agentic Optimization on Large Codebases

DGX agent

arXiv:2603.16011v2 Announce Type: replace-cross Abstract: Large language model (LLM) coding agents increasingly operate at the repository level, motivating benchmarks that evaluate their ability to op

model-releasesarxiv-cs-ai
18 May 2026
Model Releases

SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows?

DGX agent

arXiv:2605.15777v1 Announce Type: new Abstract: Computer-Using Agents (CUAs) are rapidly extending large language models (LLMs) beyond text-based reasoning toward action execution in more complex envi

model-releasesarxiv-cs-ai
18 May 2026
Model Releases

SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory

DGX agent

arXiv:2605.15710v1 Announce Type: new Abstract: Existing benchmarks for multimodal memory reasoning largely evaluate systems within pre-assembled contexts, but under-evaluate whether agents can use ev

model-releasesarxiv-cs-cl
18 May 2026
Safety

Unveiling the Black Box: A Multi-Layer Framework for Explaining Reinforcement Learning-Based Cyber Agents

DGX agent

arXiv:2505.11708v3 Announce Type: replace-cross Abstract: Reinforcement Learning (RL) agents are increasingly used to simulate sophisticated cyberattacks, but their decision-making processes remain op

safetyarxiv-cs-lg
18 May 2026
Safety

Verifiable Agentic Infrastructure: Proof-Derived Authorization for Sovereign AI Systems

DGX agent

arXiv:2605.15228v1 Announce Type: new Abstract: Modern cloud and enterprise systems rely on identity-centric authorization, assuming that callers possessing valid credentials are safe to execute comma

safetyarxiv-cs-ai
18 May 2026
Agents

Articraft: An Agentic System for Scalable Articulated 3D Asset Generation

DGX agent

arXiv:2605.15187v1 Announce Type: new Abstract: A bottleneck in learning to understand articulated 3D objects is the lack of large and diverse datasets. In this paper, we propose to leverage large lan

agentsarxiv-cs-cv
15 May 2026
Agents

CA2: Code-Aware Agent for Automated Game Testing

DGX agent

arXiv:2605.13918v1 Announce Type: cross Abstract: Automated game testing is important for verifying game functionality, but it remains a costly and time-consuming process. Manual testing often misses

agentsarxiv-cs-lg
15 May 2026
Agents

DermAgent: A Self-Reflective Agentic System for Dermatological Image Analysis with Multi-Tool Reasoning and Traceable Decision-Making

DGX agent

arXiv:2605.14403v1 Announce Type: new Abstract: Dermatological diagnosis requires integrating fine-grained visual perception with expert clinical knowledge. Although Multimodal Large Language Models (

agentsarxiv-cs-cv
15 May 2026
Model Releases

FutureSim: Replaying World Events to Evaluate Adaptive Agents

DGX agent

arXiv:2605.15188v1 Announce Type: cross Abstract: AI agents are being increasingly deployed in dynamic, open-ended environments that require adapting to new information as it arrives. To efficiently m

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory

DGX agent

arXiv:2605.15128v1 Announce Type: cross Abstract: Long-term agent memory is increasingly multimodal, yet existing evaluations rarely test whether agents preserve the visual evidence needed for later r

model-releasesarxiv-cs-cl
15 May 2026
Agents

Nexus : An Agentic Framework for Time Series Forecasting

DGX agent

arXiv:2605.14389v1 Announce Type: new Abstract: Time series forecasting is not just numerical extrapolation, but often requires reasoning with unstructured contextual data such as news or events. Whil

agentsarxiv-cs-ai
15 May 2026
Model Releases

PolitNuggets: Benchmarking Agentic Discovery of Long-Tail Political Facts

DGX agent

arXiv:2605.14002v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) embedded in agentic frameworks have transformed information retrieval from static, long context question answering into op

model-releasesarxiv-cs-ai
15 May 2026
Agents

SR-Platform: An Agentic Pipeline for Natural Language-Driven Robot Simulation Environment Synthesis

DGX agent

arXiv:2605.14700v1 Announce Type: new Abstract: Generating robot simulation environments remains a major bottleneck in simulation-based robot learning. Constructing a training-ready MuJoCo scene typic

agentsarxiv-cs-ro
15 May 2026
Safety

TERMS-Bench: Diagnosing LLM Negotiation Agents Beyond Deal Rate

DGX agent

arXiv:2605.13909v1 Announce Type: cross Abstract: Negotiation is a central mechanism of economic exchange, shaping markets, procurement, labor agreements, and resource allocation. It is also a canonic

safetyarxiv-cs-ai
15 May 2026
Model Releases

GUIGuard-Bench: Toward a General Evaluation for Privacy-Preserving GUI Agents

DGX agent

arXiv:2601.18842v3 Announce Type: replace-cross Abstract: As GUI agents increasingly rely on screenshots to perceive and operate digital environments, they may inadvertently expose sensitive informati

model-releasesarxiv-cs-ai
14 May 2026
Model Releases

MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows

DGX agent

arXiv:2603.24649v2 Announce Type: replace Abstract: Medical imaging benchmarks often evaluate VLMs on pre-selected 2D images, slices, crops, or patches, making evaluation closer to visual recognition.

model-releasesarxiv-cs-cv
14 May 2026
Agents

Moltbook Moderation: Uncovering Hidden Intent Through Multi-Turn Dialogue

DGX agent

arXiv:2605.12856v1 Announce Type: new Abstract: The emergence of multi-agent systems introduces novel moderation challenges that extend beyond content filtering. Agents with {em malicious intent} may

agentsarxiv-cs-ai
14 May 2026
Agents

Multi-Agent Systems in Emergency Departments: Validation Study on a ED Digital Twin

DGX agent

arXiv:2605.13345v1 Announce Type: new Abstract: Emergency departments (ED) face challenges in patient care and resource management. We propose to explore optimization strategies in a realistic and fle

agentsarxiv-cs-ai
14 May 2026
Agents

RDMA: Cost Effective Agent-Driven Rare Disease Mining from Electronic Health Records

DGX agent

arXiv:2507.15867v2 Announce Type: replace-cross Abstract: Rare diseases affect 1 in 10 Americans yet remain systematically underdocumented in clinical records. ICD-based systems cannot capture their b

agentsarxiv-cs-ai
14 May 2026
Model Releases

RS-Claw: Progressive Active Tool Exploration via Hierarchical Skill Trees for Remote Sensing Agents

DGX agent

arXiv:2605.13391v1 Announce Type: new Abstract: The rise of multi-modal large language models (MLLMs) is shifting remote sensing (RS) intelligence from 'see' to 'action', as OpenClaw-style frameworks

model-releasesarxiv-cs-ai
14 May 2026
Agents

What Do You Think I Think? Accounting for Human Beliefs Using Second-Order Theory of Mind

DGX agent

arXiv:2605.12745v1 Announce Type: cross Abstract: Discrepancies between an agent's actual knowledge and what a person thinks the agent knows can hinder interactions. If an agent could detect such disc

agentsarxiv-cs-ai
14 May 2026
Model Releases

Breaking Down and Building Up: Mixture of Skill-Based Vision-and-Language Navigation Agents

DGX agent

arXiv:2508.07642v3 Announce Type: replace-cross Abstract: Vision-and-Language Navigation (VLN) poses significant challenges for agents to interpret natural language instructions and navigate complex 3

model-releasesarxiv-cs-cl
13 May 2026
Agents

From Reaction to Anticipation: Proactive Failure Recovery through Agentic Task Graph for Robotic Manipulation

DGX agent

arXiv:2605.11951v1 Announce Type: new Abstract: Although robotic manipulation has made significant progress, reliable execution remains challenging because task failures are inevitable in dynamic and

agentsarxiv-cs-ro
13 May 2026
Safety

Learning Agentic Policy from Action Guidance

DGX agent

arXiv:2605.12004v1 Announce Type: new Abstract: Agentic reinforcement learning (RL) for Large Language Models (LLMs) critically depends on the exploration capability of the base policy, as training si

safetyarxiv-cs-cl
13 May 2026
Safety

SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs

DGX agent

arXiv:2605.12039v1 Announce Type: new Abstract: Skill libraries enable large language model agents to reuse experience from past interactions, but most existing libraries store skills as isolated entr

safetyarxiv-cs-cl
13 May 2026
Agents

Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery

DGX agent

arXiv:2605.08956v1 Announce Type: new Abstract: A growing body of work pursues AI scientists capable of end-to-end autonomous scientific discovery. This position paper argues that although they alread

agentsarxiv-cs-ai
12 May 2026
Model Releases

ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox

DGX agent

arXiv:2605.10787v1 Announce Type: new Abstract: Current LLM agents are proficient at calling isolated APIs but struggle with the 'last mile' of commercial software automation. In real-world scenarios,

model-releasesarxiv-cs-ai
12 May 2026
← Previous
1…7071727374…236
Next →