AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,958 results
21 May 2026

Build AI-powered dashboard automation agents with NLP on Amazon Bedrock AgentCore

IndustryDGX agent

This solution combines the power of Amazon Bedrock AgentCore, Strands Agents, and Amazon Quick transforms to deliver a secure, scalable, and intelligent system for building and operating AI agents whi

FT-Dojo: Towards Autonomous LLM Fine-Tuning with Language Agents

Model ReleasesDGX agent

arXiv:2603.01712v2 Announce Type: replace-cross Abstract: Fine-tuning large language models for vertical domains remains labor-intensive, requiring practitioners to curate data, configure training, an

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning

AgentsDGX agent

arXiv:2605.20342v1 Announce Type: new Abstract: Training large multimodal models (LMMs) via reinforcement learning (RL) to natively invoke video-processing tools (e.g., cropping) has become a promisin

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
20 May 2026

EngiAI: A Multi-Agent Framework and Benchmark Suite for LLM-Driven Engineering Design

Model ReleasesDGX agent

arXiv:2605.19743v1 Announce Type: new Abstract: Large Language Model (LLM) agents are increasingly applied to engineering design tasks, yet existing evaluation frameworks do not adequately address mul

From Intent to AI Pipelines: A Controlled Agentic Framework for Non-AI Expert Scientists

AgentsDGX agent

arXiv:2605.18764v1 Announce Type: cross Abstract: Artificial Intelligence (AI) pipelines have become integral to modern research, supporting fields such as Medical Sciences, Agriculture, and Social Sc

Measuring Safety Alignment Effects in Autonomous Security Agents

Model ReleasesDGX agent

arXiv:2605.19722v1 Announce Type: cross Abstract: Do stock safety-aligned language models and their uncensored or abliterated derivatives behave differently when run as autonomous security agents? Sin

Microsoft Senior AI developer just showed how they build AI agents with Claude at Microsoft. 34-minutes. free. By Microsoft team Opus 4.7 + …

Model ReleasesDGX agent

Microsoft Senior AI developer just showed how they build AI agents with Claude at Microsoft. 34-minutes. free. By Microsoft team Opus 4.7 + 1,400+ pre-built MCP tools plug Claude into agent → give it

Progressive Autonomy as Preference Learning: A Formalization of Trust Calibration for Agentic Tool Use

SafetyDGX agent

arXiv:2605.19151v1 Announce Type: new Abstract: We formalize trust calibration for agentic tool use (deciding when an automated agent's proposed action may execute autonomously versus require human ap

Sequential Consensus for Multi-Agent LLM Debates: A Wald-SPRT compute governor with calibration-based failure detection

Model ReleasesDGX agent

arXiv:2605.19193v1 Announce Type: new Abstract: Multi-agent LLM debate improves factuality and reasoning, but most recipes pick a fixed round count, over-spending on easy items and under-spending on h

The World Won't Stay Still: Programmable Evolution for Agent Benchmarks

Model ReleasesDGX agent

arXiv:2603.05910v2 Announce Type: replace Abstract: LLM-powered tool-calling agents fulfill user requests by interacting with environments, querying data, and invoking tools in a multi-turn process. Y

Tribal AI lands $10M in seed funding to bring metadata-native agents to the enterprise

AgentsDGX agent

Enterprise software veterans have become all too familiar with the gap between flashy artificial intelligence demos and their performance in real-world production environments. The reality is that a l

19 May 2026

Adversarial Agent Collaboration for Correctness Improvements of C to Safe Rust Translation

Model ReleasesDGX agent

arXiv:2510.03879v3 Announce Type: replace-cross Abstract: Translating C to memory-safe languages, like Rust, prevents critical memory safety vulnerabilities that are prevalent in legacy C software. Ev

AI Agents May Always Fall for Prompt Injections

SafetyDGX agent

arXiv:2605.17634v1 Announce Type: cross Abstract: Prompt injection is the most critical vulnerability in deployed AI agents. Despite recent progress, we show that the prevailing defense paradigm (data

Aurora: Unified Video Editing with a Tool-Using Agent

Local AiDGX agent

arXiv:2605.18748v1 Announce Type: new Abstract: Recent video editing models have converged on a unified conditioning design: a single diffusion transformer jointly consumes text, source video, and ref

CVE-Factory: Scaling Expert-Level Agentic Tasks for Code Security Vulnerability

Model ReleasesDGX agent

arXiv:2602.03012v2 Announce Type: replace-cross Abstract: Evaluating and improving the security capabilities of code agents requires high-quality, executable vulnerability tasks. However, existing wor

EvoMemBench: Benchmarking Agent Memory from a Self-Evolving Perspective

Model ReleasesDGX agent

arXiv:2605.18421v1 Announce Type: cross Abstract: Recent benchmarks for Large Language Model (LLM) agents mainly evaluate reasoning, planning, and execution. However, memory is also essential for agen

I/O 2026: Welcome to the agentic Gemini era

Model ReleasesDGX agent

At I/O 2026, Google announced that AI is transitioning from something users actively open to a background service that completes tasks automatically. The company introduced Gemini Spark, a new agentic

Learning from Self-Debate: Preparing Reasoning Models for Multi-Agent Debate

AgentsDGX agent

arXiv:2601.22297v2 Announce Type: replace Abstract: The reasoning abilities of large language models (LLMs) have been substantially improved by reinforcement learning with verifiable rewards (RLVR). A

MA^{2}P: A Meta-Cognitive Autonomous Intelligent Agents Framework for Complex Persuasion

AgentsDGX agent

arXiv:2605.18572v1 Announce Type: new Abstract: Persuasive dialogue generation plays a vital role in decision-making, negotiation, counseling, and behavior change, yet it remains a challenging problem

RAGA: Reading-And-Graph-building-Agent for Autonomous Knowledge Graph Construction and Retrieval-Augmented Generation

AgentsDGX agent

arXiv:2605.17072v1 Announce Type: new Abstract: Existing LLM-driven knowledge graph (KG) construction methods predominantly employ stateless batch processing pipelines, exhibiting structural deficienc

SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors

Model ReleasesDGX agent

arXiv:2605.16626v1 Announce Type: cross Abstract: Since autonomous coding agents generate complex behaviors at high-volume, we may want to use other LLMs to monitor actions to reduce the risk from dan

State Contamination in Memory-Augmented LLM Agents

SafetyDGX agent

arXiv:2605.16746v1 Announce Type: new Abstract: LLM agents increasingly rely on persistent state, including transcripts, summaries, retrieved context, and memory buffers, to support long-horizon inter

Supervising the search process produces reliable and generalizable information-seeking agents

Model ReleasesDGX agent

arXiv:2502.13957v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are transforming web search by shifting from document ranking to synthesizing answers, and are increasingly deplo

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents

Model ReleasesDGX agent

arXiv:2605.16282v1 Announce Type: cross Abstract: The rapid deployment of LLM-based autonomous agents has introduced safety risks that extend far beyond traditional LLM concerns, prompting a prolifera

The End of Trust: How Agentic AI Breaks Security Assumptions

AgentsDGX agent

arXiv:2605.16436v1 Announce Type: cross Abstract: For decades, the security of digital interaction has rested on an unacknowledged economic constraint. Attackers faced a tradeoff between the fidelity

[video] why we need a new continuity layer for long-running agents (claude did this video! all except the voice which was @elevenlabs)

Model ReleasesDGX agent

This video discusses the architectural need for a continuity layer in long-running AI agents, explaining how agents require persistent memory and state management mechanisms to maintain coherence acro

When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents

Model ReleasesDGX agent

arXiv:2601.17887v2 Announce Type: replace Abstract: Long-term memory enables large language model (LLM) agents to support personalized and sustained interactions. However, most work on personalized ag

18 May 2026

A3D: Agentic AI flow for autonomous Accelerator Design

Model ReleasesDGX agent

arXiv:2605.15237v1 Announce Type: cross Abstract: Accelerating applications through the design of hardware accelerators can significantly enhance system performance and energy efficiency. Despite adva

Argus: Evidence Assembly for Scalable Deep Research Agents

Model ReleasesDGX agent

arXiv:2605.16217v1 Announce Type: cross Abstract: Deep research agents have achieved remarkable progress on complex information seeking tasks. Even long ReAct style rollouts explore only a single traj

Coding agent tracing and evaluation: An open source tool to improve AI coding workflows

Model ReleasesDGX agent

Announcing coding harness tracing for observing, evaluating, and improving coding agent workflows across Claude Code, Cursor, Codex, GitHub Copilot, and Gemini CLI. The post Coding agent tracing and e

CryptoBench: A Dynamic Benchmark for Expert-Level Evaluation of LLM Agents in Cryptocurrency

Model ReleasesDGX agent

arXiv:2512.00417v5 Announce Type: replace Abstract: This paper introduces CryptoBench, the first expert-curated, dynamic benchmark designed to rigorously evaluate the real-world capabilities of Large

FormulaCode: Evaluating Agentic Optimization on Large Codebases

Model ReleasesDGX agent

arXiv:2603.16011v2 Announce Type: replace-cross Abstract: Large language model (LLM) coding agents increasingly operate at the repository level, motivating benchmarks that evaluate their ability to op

How do you know your document parser is ready for production? 🤔 Existing benchmarks miss what AI agents actually need. That's the gap Parse…

Model ReleasesDGX agent

How do you know your document parser is ready for production? 🤔 Existing benchmarks miss what AI agents actually need. That's the gap ParseBench, the first doc OCR benchmark for AI agents, fills. We'l

SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows?

Model ReleasesDGX agent

arXiv:2605.15777v1 Announce Type: new Abstract: Computer-Using Agents (CUAs) are rapidly extending large language models (LLMs) beyond text-based reasoning toward action execution in more complex envi

SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory

Model ReleasesDGX agent

arXiv:2605.15710v1 Announce Type: new Abstract: Existing benchmarks for multimodal memory reasoning largely evaluate systems within pre-assembled contexts, but under-evaluate whether agents can use ev

Today in AI Engineering (May 17) • Nous Research ships Hermes Agent v0.14.0: Grok subs, Codex runtime, Windows beta • LangSmith Engine relea…

Model ReleasesDGX agent

Today in AI Engineering (May 17) • Nous Research ships Hermes Agent v0.14.0: Grok subs, Codex runtime, Windows beta • LangSmith Engine releases trace issue clustering, drafts PRs and evals from prod t

Unveiling the Black Box: A Multi-Layer Framework for Explaining Reinforcement Learning-Based Cyber Agents

SafetyDGX agent

arXiv:2505.11708v3 Announce Type: replace-cross Abstract: Reinforcement Learning (RL) agents are increasingly used to simulate sophisticated cyberattacks, but their decision-making processes remain op

Verifiable Agentic Infrastructure: Proof-Derived Authorization for Sovereign AI Systems

SafetyDGX agent

arXiv:2605.15228v1 Announce Type: new Abstract: Modern cloud and enterprise systems rely on identity-centric authorization, assuming that callers possessing valid credentials are safe to execute comma

15 May 2026

Articraft: An Agentic System for Scalable Articulated 3D Asset Generation

AgentsDGX agent

arXiv:2605.15187v1 Announce Type: new Abstract: A bottleneck in learning to understand articulated 3D objects is the lack of large and diverse datasets. In this paper, we propose to leverage large lan

CA2: Code-Aware Agent for Automated Game Testing

AgentsDGX agent

arXiv:2605.13918v1 Announce Type: cross Abstract: Automated game testing is important for verifying game functionality, but it remains a costly and time-consuming process. Manual testing often misses

DermAgent: A Self-Reflective Agentic System for Dermatological Image Analysis with Multi-Tool Reasoning and Traceable Decision-Making

AgentsDGX agent

arXiv:2605.14403v1 Announce Type: new Abstract: Dermatological diagnosis requires integrating fine-grained visual perception with expert clinical knowledge. Although Multimodal Large Language Models (

FutureSim: Replaying World Events to Evaluate Adaptive Agents

Model ReleasesDGX agent

arXiv:2605.15188v1 Announce Type: cross Abstract: AI agents are being increasingly deployed in dynamic, open-ended environments that require adapting to new information as it arrives. To efficiently m

It has been a pleasure collaborating with the @NVIDIAAI team to ensure that Hermes Agent runs perfectly on DGX Spark!

Local AiDGX agent

It has been a pleasure collaborating with the @NVIDIAAI team to ensure that Hermes Agent runs perfectly on DGX Spark! Run @NousResearch's Hermes Agent fully locally on DGX Spark. 🚀 Our newest playbook

MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory

Model ReleasesDGX agent

arXiv:2605.15128v1 Announce Type: cross Abstract: Long-term agent memory is increasingly multimodal, yet existing evaluations rarely test whether agents preserve the visual evidence needed for later r

Nexus : An Agentic Framework for Time Series Forecasting

AgentsDGX agent

arXiv:2605.14389v1 Announce Type: new Abstract: Time series forecasting is not just numerical extrapolation, but often requires reasoning with unstructured contextual data such as news or events. Whil

PolitNuggets: Benchmarking Agentic Discovery of Long-Tail Political Facts

Model ReleasesDGX agent

arXiv:2605.14002v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) embedded in agentic frameworks have transformed information retrieval from static, long context question answering into op

SR-Platform: An Agentic Pipeline for Natural Language-Driven Robot Simulation Environment Synthesis

AgentsDGX agent

arXiv:2605.14700v1 Announce Type: new Abstract: Generating robot simulation environments remains a major bottleneck in simulation-based robot learning. Constructing a training-ready MuJoCo scene typic

TERMS-Bench: Diagnosing LLM Negotiation Agents Beyond Deal Rate

SafetyDGX agent

arXiv:2605.13909v1 Announce Type: cross Abstract: Negotiation is a central mechanism of economic exchange, shaping markets, procurement, labor agreements, and resource allocation. It is also a canonic

14 May 2026

Continual Harness: Online Adaptation for Self-Improving Foundation Agents [R]

ResearchDGX agent

Continual Harness proposes an approach to online adaptation for foundation agents that moves beyond traditional gradient-based retraining by introducing a dual-agent architecture (Teacher/Student) wit

GUIGuard-Bench: Toward a General Evaluation for Privacy-Preserving GUI Agents

Model ReleasesDGX agent

arXiv:2601.18842v3 Announce Type: replace-cross Abstract: As GUI agents increasingly rely on screenshots to perceive and operate digital environments, they may inadvertently expose sensitive informati

Kimi K2.6 is now open-weight #1 on Finance Agent Benchmark V2.

Model ReleasesDGX agent

Kimi K2.6 is now open-weight #1 on Finance Agent Benchmark V2. Can AI do the job of a financial analyst? We just released V2 of our Finance Agent Benchmark and tested the frontier models. The results

MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows

Model ReleasesDGX agent

arXiv:2603.24649v2 Announce Type: replace Abstract: Medical imaging benchmarks often evaluate VLMs on pre-selected 2D images, slices, crops, or patches, making evaluation closer to visual recognition.

Moltbook Moderation: Uncovering Hidden Intent Through Multi-Turn Dialogue

AgentsDGX agent

arXiv:2605.12856v1 Announce Type: new Abstract: The emergence of multi-agent systems introduces novel moderation challenges that extend beyond content filtering. Agents with {em malicious intent} may

Multi-Agent Systems in Emergency Departments: Validation Study on a ED Digital Twin

AgentsDGX agent

arXiv:2605.13345v1 Announce Type: new Abstract: Emergency departments (ED) face challenges in patient care and resource management. We propose to explore optimization strategies in a realistic and fle

RDMA: Cost Effective Agent-Driven Rare Disease Mining from Electronic Health Records

AgentsDGX agent

arXiv:2507.15867v2 Announce Type: replace-cross Abstract: Rare diseases affect 1 in 10 Americans yet remain systematically underdocumented in clinical records. ICD-based systems cannot capture their b

RS-Claw: Progressive Active Tool Exploration via Hierarchical Skill Trees for Remote Sensing Agents

Model ReleasesDGX agent

arXiv:2605.13391v1 Announce Type: new Abstract: The rise of multi-modal large language models (MLLMs) is shifting remote sensing (RS) intelligence from 'see' to 'action', as OpenClaw-style frameworks

The Building AI Agents with MongoDB and LangGraph Skill Badge can be yours on May 28th. Here's how you earn it 👉 Join our LIVE workshop, bu…

Model ReleasesDGX agent

The Building AI Agents with MongoDB and LangGraph Skill Badge can be yours on May 28th. Here's how you earn it 👉 Join our LIVE workshop, build a real agent with MongoDB, Claude Sonnet, and LangGraph,

What Do You Think I Think? Accounting for Human Beliefs Using Second-Order Theory of Mind

AgentsDGX agent

arXiv:2605.12745v1 Announce Type: cross Abstract: Discrepancies between an agent's actual knowledge and what a person thinks the agent knows can hinder interactions. If an agent could detect such disc

13 May 2026

Breaking Down and Building Up: Mixture of Skill-Based Vision-and-Language Navigation Agents

Model ReleasesDGX agent

arXiv:2508.07642v3 Announce Type: replace-cross Abstract: Vision-and-Language Navigation (VLN) poses significant challenges for agents to interpret natural language instructions and navigate complex 3

From Reaction to Anticipation: Proactive Failure Recovery through Agentic Task Graph for Robotic Manipulation

AgentsDGX agent

arXiv:2605.11951v1 Announce Type: new Abstract: Although robotic manipulation has made significant progress, reliable execution remains challenging because task failures are inevitable in dynamic and

← Previous
1…96979899100…300
Next →