AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,959 results
15 May 2026

Holistic Evaluation and Failure Diagnosis of AI Agents

Model ReleasesDGX agent

arXiv:2605.14865v1 Announce Type: new Abstract: AI agents execute complex multi-step processes, but current evaluation falls short: outcome metrics report success or failure without explaining why, an

Hypergraph Enterprise Agentic Reasoner over Heterogeneous Business Systems

AgentsDGX agent

arXiv:2605.14259v1 Announce Type: new Abstract: Applying Large Language Models (LLMs) to heterogeneous enterprise systems is hindered by hallucinations and failures in multi-hop, n-ary reasoning. Exis

LEMON: Learning Executable Multi-Agent Orchestration via Counterfactual Reinforcement Learning

Local AiDGX agent

arXiv:2605.14483v1 Announce Type: new Abstract: Large language models (LLMs) have become a strong foundation for multi-agent systems, but their effectiveness depends heavily on orchestration design. A

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

MAPLE: Latent Multi-Agent Play for End-to-End Autonomous Driving

SafetyDGX agent

arXiv:2605.14201v1 Announce Type: cross Abstract: Vision-language-action (VLA) models are effective as end-to-end motion planners, but can be brittle when evaluated in closed-loop settings due to bein

OpenAI keeps shuffling its executives in bid to win AI agent battle

AgentsDGX agent

OpenAI announced yet another reorganization Friday, consolidating certain areas and making company president Greg Brockman the official lead of all things product. In a memo viewed by The Verge, Brock

pi-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows

Model ReleasesDGX agent

arXiv:2605.14678v1 Announce Type: new Abstract: The rise of personal assistant agents, e.g., OpenClaw, highlights the growing potential of large language models to support users across everyday life a

Self-Distilled Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2605.15155v1 Announce Type: cross Abstract: Reinforcement learning (RL) has emerged as a central paradigm for post-training LLM agents, yet its trajectory-level reward signal provides only coars

14 May 2026

Anthropic announces ‘programmatic credit pool’ as agentic tool use rises

Model ReleasesDGX agent

Anthropic PBC, the developer and provider of the Claude artificial intelligence model family, said it’s offering a special credit pool for users who want to use agentic tools with its large language m

Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents

ResearchDGX agent

arXiv:2605.12894v1 Announce Type: new Abstract: Large Language Model (LLM) agents are increasingly deployed in settings where they interact with a wide variety of people, including users who are uncle

Cloud CISO Perspectives: How Google + Wiz changes multicloud strategy for CISOs

Model ReleasesDGX agent

Welcome to the first Cloud CISO Perspectives for May 2026. Today, Vinod D’Souza, director, Office of the CISO, shares highlights from his RSA Conference fireside chat with Anthony Belfiore, chief stra

Data readiness for agentic AI in financial services

AgentsDGX agent

Financial services companies have unique needs when it comes to business AI. They operate in one of the most highly regulated sectors while responding to external events that are updated by the second

MARLIN: Multi-Agent Game-Theoretic Reinforcement Learning for Sustainable LLM Inference in Cloud Datacenters

AgentsDGX agent

arXiv:2605.13496v1 Announce Type: cross Abstract: Large Language Models (LLMs) have become increasingly prevalent in cloud-based platforms, propelled by the introduction of AI-based consumer and enter

Mechanism Plausibility in Generative Agent-Based Modeling

AgentsDGX agent

arXiv:2605.12824v1 Announce Type: cross Abstract: Large language models (LLMs) can generate high-level diverse phenomena without explicitly programmed rules. This capability has led to their adoption

Multi-Dimensional Behavioral Evaluation of Agentic Stock Prediction Systems Using Large Language Model Judges with Closed-Loop Reinforcement Learning Feedback

AgentsDGX agent

arXiv:2605.05739v2 Announce Type: replace-cross Abstract: Forecast evaluation in finance has relied on aggregate accuracy metrics and predictive-accuracy tests built on point-forecast errors. These in

Plan Before You Trade: Inference-Time Optimization for RL Trading Agents

Model ReleasesDGX agent

arXiv:2605.12653v1 Announce Type: cross Abstract: Reinforcement learning agents for portfolio management are typically trained and deployed as static policies, with no mechanism for using price foreca

ScioMind: Cognitively Grounded Multi-Agent Social Simulation with Anchoring-Based Belief Dynamics and Dynamic Profiles

SafetyDGX agent

arXiv:2605.13725v1 Announce Type: new Abstract: Large language model (LLM)-based multi-agent simulation offers a powerful testbed for studying social opinion dynamics. Yet current approaches often ado

Think Twice, Act Once: Verifier-Guided Action Selection For Embodied Agents

SafetyDGX agent

arXiv:2605.12620v1 Announce Type: new Abstract: Building generalist embodied agents capable of solving complex real-world tasks remains a fundamental challenge in AI. Multimodal Large Language Models

Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2605.13119v1 Announce Type: cross Abstract: Vision-language-action (VLA) models are effective robot action executors, but they remain limited on long-horizon tasks due to the dual burden of exte

13 May 2026

Agent-BRACE: Decoupling Beliefs from Actions in Long-Horizon Tasks via Verbalized State Uncertainty

Model ReleasesDGX agent

arXiv:2605.11436v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed on long-horizon tasks in partially observable environments, where they must act while inferring a

Boomi CEO: The agents are ready, but the enterprise data foundation beneath them might not be

AgentsDGX agent

Somewhere inside every large company, there is a long list of AI pilots — promising, applauded, then quietly shelved. The culprit, more often than not, is a data activation failure. The gap between AI

GRAFT-ATHENA: Self-Improving Agentic Teams for Autonomous Discovery and Evolutionary Numerical Algorithms

Model ReleasesDGX agent

arXiv:2605.11117v1 Announce Type: new Abstract: Scientific discovery can be modeled as a sequence of probabilistic decisions that map physical problems to numerical solutions. Recent agentic AI system

How we use Alyx to build Alyx: How to build an AI agent feedback loop

AgentsDGX agent

How Arize uses Alyx to debug Alyx: searching dense traces, aggregating failures, triaging dogfooding issues, and closing the AI engineering feedback loop. The post How we use Alyx to build Alyx: How t

@nvidia Nemotron native support in Deep Agents 0.6 #interrupt #langchain

Model ReleasesDGX agent

NVIDIA Nemotron models received native support integration in Deep Agents version 0.6, enabling improved language model capabilities within the LangChain framework. This update allows developers to le

Predictive Maps of Multi-Agent Reasoning: A Successor-Representation Spectrum for LLM Communication Topologies

SafetyDGX agent

arXiv:2605.11453v1 Announce Type: cross Abstract: Practitioners deploying multi-agent large language model (LLM) systems must currently choose between communication topologies such as chain, star, mes

Probabilistic Modeling of Latent Agentic Substructures in Deep Neural Networks

SafetyDGX agent

arXiv:2509.06701v2 Announce Type: replace Abstract: We develop a theory of intelligent agency grounded in probabilistic modeling for neural models. Agents are represented as outcome distributions with

Transform Video Into Instantly Searchable, Actionable Intelligence with AI Agents and Skills

HardwareDGX agent

NVIDIA's video analytics AI agents analyze and process large volumes of video data through natural language tasks to provide critical insights , powered by vision language models, large language model

12 May 2026

A Reflective Storytelling Agent for Older Adults: Integrating Argumentation Schemes and Argument Mining in LLM-Based Personalised Narratives

AgentsDGX agent

arXiv:2605.10531v1 Announce Type: new Abstract: This work investigates whether knowledge-driven large language model (LLM)-based storytelling can support purposeful narrative interaction with a digita

A Versatile AI Agent for Rare Disease Diagnosis and Risk Gene Prioritization

AgentsDGX agent

arXiv:2605.06226v2 Announce Type: replace Abstract: Accurate and timely diagnosis is essential for effective treatment, particularly in the context of rare diseases. However, current diagnostic workfl

Agentic Performance at the Edge: Insights from Benchmarking

Model ReleasesDGX agent

arXiv:2605.10384v1 Announce Type: new Abstract: Agentic artificial intelligence (AI) is a natural fit for Internet of Things (IoT) and edge systems, but edge deployments are often constrained to model

An agentic framework for gravitational-wave counterpart association in the multi-messenger era

AgentsDGX agent

arXiv:2605.10584v1 Announce Type: cross Abstract: With the detection of gravitational waves (GWs), multi-messenger astronomy has opened a new window for advancing our understanding of astrophysics, de

Can Deep Research Agents Retrieve and Organize? Evaluating the Synthesis Gap with Expert Taxonomies

Model ReleasesDGX agent

arXiv:2601.12369v3 Announce Type: replace Abstract: Deep Research Agents increasingly automate survey generation, yet whether they match human experts at retrieving essential papers and organizing the

Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces

Model ReleasesDGX agent

arXiv:2605.09497v1 Announce Type: new Abstract: Vision-language model (VLM) based web agents demonstrate impressive autonomous GUI interaction but remain vulnerable to deceptive interface elements. Ex

Dynamic Skill Lifecycle Management for Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2605.10923v1 Announce Type: cross Abstract: Large language model agents increasingly rely on external skills to solve complex tasks, where skills act as modular units that extend their capabilit

EconWebArena: Benchmarking Autonomous Agents on Economic Tasks in Realistic Web Environments

Model ReleasesDGX agent

arXiv:2506.08136v3 Announce Type: replace Abstract: We introduce EconWebArena, a benchmark for evaluating autonomous agents on complex, multimodal economic tasks in realistic web environments. The ben

Exclusive: ClickUp endows its Brain assistant with agentic capabilities

AgentsDGX agent

ClickUp, the business name of Mango Technologies Inc., today is introducing what it says is a major overhaul of its Brain artificial intelligence assisting inside its workplace collaboration platform,

FORTIS: Benchmarking Over-Privilege in Agent Skills

Model ReleasesDGX agent

arXiv:2605.09163v1 Announce Type: new Abstract: Large language model agents increasingly operate through an intermediate skill layer that mediates between user intent and concrete task execution. This

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World

ApplicationsDGX agent

arXiv:2605.10834v1 Announce Type: new Abstract: AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide limited guidance on which will perfor

GenCellAgent: Generalizable, Training-Free Cellular Image Segmentation via Large Language Model Agents

AgentsDGX agent

arXiv:2510.13896v2 Announce Type: replace-cross Abstract: Cellular image segmentation is essential for quantitative biology yet remains difficult due to heterogeneous modalities, morphological variabi

MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring

Model ReleasesDGX agent

arXiv:2605.09684v1 Announce Type: cross Abstract: We introduce a red-teaming methodology that exposes harder-to-catch attacks for coding-agent monitors, suggesting that current practices may under-eli

MOTOR-Bench: A Real-world Dataset and Multi-agent Framework for Zero-shot Human Mental State Understanding

Model ReleasesDGX agent

arXiv:2605.09703v1 Announce Type: new Abstract: Understanding human mental states from natural behavior is crucial for intelligent systems in the real world. However, most current research focuses on

PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools

SafetyDGX agent

arXiv:2604.01532v2 Announce Type: replace Abstract: LLM agents are beginning to invoke industrial asset-management tools through the Model Context Protocol (MCP), yet whether they can act reliably on

Red Hat expands agentic AI strategy with new inference, automation and sovereignty capabilities

AgentsDGX agent

IBM Corp. subsidiary Red Hat today is unveiling a broad set of product and partnership announcements aimed at helping enterprises put artificial intelligence into operation, modernize infrastructure a

Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient?

Model ReleasesDGX agent

arXiv:2605.10848v1 Announce Type: cross Abstract: Does a lexical retriever suffice as large language models (LLMs) become more capable in an agentic loop? This question naturally arises when building

SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution

Model ReleasesDGX agent

arXiv:2605.08366v1 Announce Type: new Abstract: We introduce SWE Atlas, a benchmark suite for coding agents spanning three professional software engineering workflows: Codebase Q&A (124 tasks), Test W

Trustworthy AI: Ensuring Reliability and Accountability from Models to Agents

SafetyDGX agent

arXiv:2605.08964v1 Announce Type: new Abstract: In this thesis, we develop algorithms with theoretical guarantees for ensuring reliability and accountability of Machine Learning (ML) systems. As ML sy

UPA: Unsupervised Prompt Agent via Tree-Based Search and Selection

Local AiDGX agent

arXiv:2601.23273v2 Announce Type: replace Abstract: Prompt agents have recently emerged as a promising paradigm for automated prompt optimization, framing prompt discovery as a sequential decision-mak

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models

AgentsDGX agent

arXiv:2605.10106v1 Announce Type: cross Abstract: Recent advances in Multi-modal Large Language Models (MLLMs) target 3D spatial intelligence, yet the progress has been largely driven by post-training

Voice choice shapes how an agent feels to users, from fintech support to healthcare intake to entertainment. Try voice finder: https://findt…

AgentsDGX agent

Voice selection significantly impacts user perception and experience across various applications, including financial services support, healthcare intake processes, and entertainment platforms. Togeth

11 May 2026

ARMOR: An Agentic Framework for Reaction Feasibility Prediction via Adaptive Utility-aware Multi-tool Reasoning

AgentsDGX agent

arXiv:2605.07103v1 Announce Type: new Abstract: Reaction feasibility prediction, as a fundamental problem in computational chemistry, has benefited from diverse tools enabled by recent advances in art

Can David Beat Goliath? On Multi-Hop Reasoning with Resource-Constrained Agents

SafetyDGX agent

arXiv:2601.21699v2 Announce Type: replace Abstract: Multi-turn reasoning agents solve complex questions by decomposing them into intermediate retrieval or tool-use steps, for accumulating supporting e

Contextual Invertible World Models: A Neuro-Symbolic Agentic Framework for Colorectal Cancer Drug Response

AgentsDGX agent

arXiv:2603.02274v2 Announce Type: replace-cross Abstract: Precision oncology is currently limited by the small-N, large-P paradox, where high-dimensional genomic data is abundant but pharmacological r

LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling

AgentsDGX agent

arXiv:2605.08083v1 Announce Type: new Abstract: Test-time scaling (TTS) has become an effective approach for improving large language model performance by allocating additional computation during infe

MedExAgent: Training LLM Agents to Ask, Examine, and Diagnose in Noisy Clinical Environments

AgentsDGX agent

arXiv:2605.07058v1 Announce Type: cross Abstract: Real-world clinical diagnosis is a complex process in which the doctor is required to obtain information from both interaction with the patient and co

OASES: Outcome-Aligned Search-Evaluation Co-Training for Agentic Search

SafetyDGX agent

arXiv:2604.03675v2 Announce Type: replace Abstract: Agentic search enables language models to solve knowledge-intensive tasks by adaptively acquiring external evidence over multiple steps. Reinforceme

Quiq extends its AI agent platform into voice as enterprise rollouts move past pilots

AgentsDGX agent

Artificial intelligence customer service platform provider Quiq Inc. today launched a new voice product and refreshed its brand, betting that customer experience teams are ready to move beyond isolate

RelAgent: LLM Agents as Data Scientists for Relational Learning

AgentsDGX agent

arXiv:2605.07840v1 Announce Type: new Abstract: Relational learning is a challenging problem that has motivated a wide range of approaches, including graph-based models (e.g., graph neural networks, g

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence

SafetyDGX agent

arXiv:2605.06230v2 Announce Type: replace Abstract: As large models evolve from conversational assistants into autonomous agents, challenges increasingly arise from long-horizon decision making, tool

When Routine Chats Turn Toxic: Unintended Long-Term State Poisoning in Personalized Agents

Model ReleasesDGX agent

arXiv:2605.06731v1 Announce Type: cross Abstract: Personalized LLM agents maintain persistent cross-session state to support long-horizon collaboration. Yet, this persistence introduces a subtle but c

8 May 2026

I built an autonomous agent that lives inside her own source code— 7 days, 480 commits, multi-provider (DeepSeek / ChatGPT / Ollama)

Model ReleasesDGX agent

This post describes a project where the developer created an autonomous AI agent capable of modifying and executing its own source code across a 7-day development period, integrating multiple language

7 May 2026

A Skill-Based AI Agentic Pipeline for Library of Congress Subject Indexing

SafetyDGX agent

arXiv:2605.03537v1 Announce Type: cross Abstract: This paper presents a modular AI agentic skill pipeline for automating subject indexing with Library of Congress Subject Headings (LCSH). Subject inde

← Previous
1…108109110111112…300
Next →