AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,964 results
31 May 2026

/goal and other fully automated AI agents are cool, but not a great model for the future of work with people. Instead you want your AI to kn…

ApplicationsDGX agent

/goal and other fully automated AI agents are cool, but not a great model for the future of work with people. Instead you want your AI to know when to ask you GOOD questions, maybe because it is stuck

We need more coding and agent traces public sharing to build datasets and better open source models! Lots of people contributing already, yo…

Model ReleasesDGX agent

We need more coding and agent traces public sharing to build datasets and better open source models! Lots of people contributing already, you should share yours too! https://huggingface.co/datasets?se

29 May 2026

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
SafetyDGX agent

arXiv:2605.29697v1 Announce Type: new Abstract: In Agentic Search, trajectory-level outcome rewards fail to quantify the behavioral contributions of individual steps, while existing step-level reward

Deep Agents v0.6 makes harness profiles a first-class abstraction. Now, you can get production-grade performance from models like @Kimi_Moon…

ApplicationsDGX agent

Deep Agents v0.6 makes harness profiles a first-class abstraction. Now, you can get production-grade performance from models like @Kimi_Moonshot, @Alibaba_Qwen, and @DeepSeek_ai at 20x+ lower cost tha

GAPD: Gold-Action Policy Distillation for Agentic Reinforcement Learning in Knowledge Base Question Answering

SafetyDGX agent

arXiv:2605.29584v1 Announce Type: new Abstract: Reinforcement learning (RL) is a natural fit for agentic knowledge base question answering (KBQA), where a model must issue executable actions, observe

GroundAct: Can LLM Agents Ground Actions in Environmental States?

Model ReleasesDGX agent

arXiv:2508.05614v2 Announce Type: replace-cross Abstract: LLM agents achieve 85-96% success on tasks where instructions fully specify the action, but drop to 29-53% when action feasibility depends on

I watched the LangChain keynote expecting product releases. Got something better — the clearest breakdown I've heard of why most agents neve…

ApplicationsDGX agent

I watched the LangChain keynote expecting product releases. Got something better — the clearest breakdown I've heard of why most agents never make it past demo. One thesis lands hard. Here's what sepa

InsightEval: An Expert-Curated Benchmark for Assessing Insight Discovery in LLM-Driven Data Agents

Model ReleasesDGX agent

arXiv:2511.22884v2 Announce Type: replace Abstract: Data analysis has become an indispensable part of scientific research. To discover the latent knowledge and insights hidden within massive datasets,

Locally Coherent, Globally Incoherent: Bounding Compositional Incoherence in Multi-Component LLM Agents

Local AiDGX agent

arXiv:2605.30335v1 Announce Type: new Abstract: Multi-component LLM agents assemble probabilistic claims from components that each see only part of a joint problem; the composition can violate basic p

Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents

SafetyDGX agent

arXiv:2605.29447v1 Announce Type: cross Abstract: While GUI agents have advanced rapidly, they often lack the robustness to recover from their own errors, hindering real-world deployment. To bridge th

Salesforce published a detailed writeup on going agentic with Claude Code. A couple things jumped out. A migration they'd scoped at 231 days…

Model ReleasesDGX agent

Salesforce published a detailed writeup on going agentic with Claude Code. A couple things jumped out. A migration they'd scoped at 231 days shipped in 13. One PR delivered 21 endpoints at 100% test c

Scaling Laws for Agent Harnesses via Effective Feedback Compute

Model ReleasesDGX agent

arXiv:2605.29682v1 Announce Type: new Abstract: Agent harnesses increasingly determine the performance of language-model systems by deciding how models call tools, receive feedback, verify intermediat

SURGENT: A Surgical Multi-Agent Assistance System Across the Perioperative Workflow

Model ReleasesDGX agent

arXiv:2605.29368v1 Announce Type: cross Abstract: The intricate nature of modern surgical care necessitates intelligent systems that can synthesize extensive patient records, support collaborative dec

28 May 2026

As agents use more context, input tokens have become the majority of price-equivalent token costs.

ToolsDGX agent

As AI agents handle increasingly complex tasks with extended context windows, the cost structure of LLM usage has shifted such that input tokens now represent the majority of price-equivalent expenses

Beyond One Path: Evaluating and Enhancing Divergent Thinking in Interactive LLM Agents

Model ReleasesDGX agent

arXiv:2605.28465v1 Announce Type: new Abstract: Divergent thinking is a core dimension of creativity, yet existing evaluations of Large Language Models (LLMs) treat them as single-turn text generation

ClinicalAgents: Multi-Agent Orchestration for Clinical Decision Making with Dual-Memory

Model ReleasesDGX agent

arXiv:2603.26182v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have demonstrated potential in healthcare, they often struggle with the complex, non-linear reasoning required fo

CPPO: Contrastive Perception Policy Optimization for VLM Agents

SafetyDGX agent

arXiv:2601.00501v2 Announce Type: replace Abstract: We introduce CPPO, a Contrastive Perception Policy Optimization method for finetuning vision--language models (VLMs). Reliable perception is a core

Diagnosing Live Within-Policy Instruction Conflicts in LLM Agents with Witnessed Resolution Profiles

SafetyDGX agent

arXiv:2605.27784v1 Announce Type: new Abstract: LLM agents are governed by long-lived natural-language prompt policies, but individually reasonable standing rules can interact in uninspected ways. We

Evaluating the Realism of LLM-powered Social Agents: A Case Study of Reactions to Spanish Online News

SafetyDGX agent

arXiv:2605.28598v1 Announce Type: cross Abstract: LLM-powered social agents are increasingly used to simulate online social behavior, yet their realism remains difficult to validate. Existing work has

Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills

Model ReleasesDGX agent

arXiv:2604.05333v3 Announce Type: replace Abstract: Modern LLM agents increasingly rely on reusable skills, and as they interact with personal applications, web browsers, and other interfaces, skill l

ICAN-Deploy: Identity-Stable Canary Deployment for Safety-Critical Embodied Agents

SafetyDGX agent

arXiv:2605.28097v1 Announce Type: new Abstract: Canary deployment routes a fraction of traffic to a new software version, monitors metrics, and rolls back on regression. Mainstream controllers (Argo R

Identifying and Mitigating Bottlenecks in Role-Playing Agents: A Systematic Study of Disentangling Character Profile Axes

SafetyDGX agent

arXiv:2601.04716v3 Announce Type: replace Abstract: While Large Language Model (LLM) role-playing agents have advanced rapidly, it remains unclear which profile elements genuinely drive role-playing q

LCO: LLM-based Constraint Optimization for Safer Agentic LLMs in Real-world Tasks

Model ReleasesDGX agent

arXiv:2605.27375v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly acting as autonomous agents, but their continuous interaction with the environment can lead to in-context

Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning

Model ReleasesDGX agent

arXiv:2605.27960v1 Announce Type: new Abstract: Despite their popularity and success, Multimodal Large Language Models (MLLMs) often struggle to interpret images accurately, which limits their reasoni

SHIPPED. Mistral Vibe is now the AI agent for long-horizon productivity and coding, and the home for Work mode, Code mode, the CLI, and a br…

Model ReleasesDGX agent

Mistral AI has released Mistral Vibe, an AI agent designed for long-horizon productivity and coding tasks, featuring Work mode, Code mode, a CLI, and additional capabilities. The product consolidates

SKILLC: Learning Autonomous Skill Internalization in LLM Agents via Contrastive Credit Assignment

SafetyDGX agent

arXiv:2605.27899v1 Announce Type: new Abstract: Structured skill prompts improve exploration in long-horizon agentic reinforcement learning (RL). Skill-augmented RL methods retain external skills at i

27 May 2026

A deep dive into how Anthropic's Claude Code and Peter Steinberger's OpenClaw unleashed the AI agent revolution that is rapidly transforming modern computing (Steven Levy/Wired)

Model ReleasesDGX agent

Steven Levy / Wired: A deep dive into how Anthropic's Claude Code and Peter Steinberger's OpenClaw unleashed the AI agent revolution that is rapidly transforming modern computing — The definitive stor

AGORA: Adapter-Grounded Observation-Action Retention for Inference-Free Prompt Compression in LLM Agents

Model ReleasesDGX agent

arXiv:2605.26596v1 Announce Type: new Abstract: The token-level extractive compressors widely used for general LM context are structurally inappropriate for LLM agents: across 17 (env, backbone, metho

AlignEvoSkill: Towards Knowledge-Aware and Task-Aligned Agent Skill Evolution

SafetyDGX agent

arXiv:2506.23149v2 Announce Type: replace Abstract: Reusable skills play a key role in improving LLM-based agents, but existing skill-evolution methods often fail to ensure that evolved skills both co

Building AI agents for business support using Amazon Bedrock AgentCore

IndustryDGX agent

In this post, we share how the AWS Generative AI Innovation Center (GenAIIC) collaborated with Works Human Intelligence (WHI) to build two AI agents using Amazon Bedrock AgentCore. We discuss the chal

Foundations of a Time-Consistent Counterfactual Actuarial Runtime for Autonomous AI Agents

SafetyDGX agent

arXiv:2605.26508v1 Announce Type: cross Abstract: We propose a foundational runtime actuarial layer for autonomous AI agents in which every side-effect-bearing action carries a time-consistent, counte

From Norms to Indicators (N2I-RAG): An Agentic Retrieval-Augmented Generation Framework for Legal Indicator Computation

SafetyDGX agent

arXiv:2605.26926v1 Announce Type: new Abstract: Computing legal indicators from normative texts is a key task in legal monitoring and policy evaluation, but presents significant challenges due to the

How AI is starting to dismantle the hegemony of the Big Four consultancies and other large firms, as AI agents help smaller consultancies handle big workloads (Financial Times)

IndustryDGX agent

Financial Times: How AI is starting to dismantle the hegemony of the Big Four consultancies and other large firms, as AI agents help smaller consultancies handle big workloads — The technology opens t

In many cases, I stopped answering my agents. Instead, I give them my criteria and let them answer

ToolsDGX agent

A manager describes shifting their leadership approach by providing agents (team members) with decision-making criteria rather than directly answering their questions, enabling them to develop problem

It's Not the Capability: Harness Sensitivity Is Non-Monotone Across LLM Agent Tiers

Model ReleasesDGX agent

arXiv:2605.26731v1 Announce Type: new Abstract: A prevalent assumption in LLM agent deployment holds that more structured harnesses universally improve reliability, and that higher-capability models n

Multi-Agent Causal Discovery Using Large Language Models

Model ReleasesDGX agent

arXiv:2407.15073v4 Announce Type: replace Abstract: Causal discovery aims to identify causal relationships between variables and is a fundamental problem across the sciences. Traditional statistical c

Natural Language Query to Configuration for Retrieval Agents

ResearchDGX agent

arXiv:2605.27361v1 Announce Type: new Abstract: Modern retrieval agents expose many configuration choices -- LLM, retriever, number of documents, number of hops, and synthesis strategy -- each shaping

New skill from K-Dense: LiteParse in Scientific Agent Skills — built for fast, local research paper ingestion. Your AI co-scientist can now:…

Local AiDGX agent

New skill from K-Dense: LiteParse in Scientific Agent Skills — built for fast, local research paper ingestion. Your AI co-scientist can now: * Parse PDFs and supplementary files on your machine (no do

Towards Feedback-to-Plan Decisions for Self-Evolving LLM Agents in CUDA Kernel Generation

HardwareDGX agent

arXiv:2605.26720v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong empirical gains as self-evolving agents for CUDA kernel generation, driven by feedback-conditioned planni

26 May 2026

Automated Benchmark Auditing for AI Agents and Large Language Models

Model ReleasesDGX agent

arXiv:2605.26079v1 Announce Type: new Abstract: Modern AI benchmarks operate at a complexity that outpaces traditional verification methods. Tasks authored by domain experts often contain implicit ass

Code2UML: Agentic LLMs with context engineering for scalable software visualization

Model ReleasesDGX agent

arXiv:2605.24453v1 Announce Type: cross Abstract: Large Language Model (LLM)-based code analysis tools are adopted to automate software documentation tasks. However, the scalability of these approache

ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions

Model ReleasesDGX agent

arXiv:2605.24279v1 Announce Type: new Abstract: A frontier language model's acknowledged 'helpful programming assistant' persona does not survive long agentic-coding sessions in the deployment regime

Contractual Skills: A GovernSpec Design Framework for Enterprise AI Agents

SafetyDGX agent

arXiv:2605.22634v2 Announce Type: replace-cross Abstract: Skills have become a practical packaging mechanism for agent instructions, workflows, scripts, and reference materials. In enterprise settings

everyone is talking about self-optimizing loops in software & agents. but what does that actually mean? in my mind, it's a system that obser…

TutorialsDGX agent

everyone is talking about self-optimizing loops in software & agents. but what does that actually mean? in my mind, it's a system that observes it's own outputs, evaluates them, and uses that signal t

From Facts to Insights: A Persona-Driven Dual Memory Framework and Dataset for Role-Playing Agents

Model ReleasesDGX agent

arXiv:2605.25693v1 Announce Type: new Abstract: While role-playing agents excel in short-term interactions, long-term conversations overwhelm context windows, motivating external memory frameworks. Cu

Introducing CHI-Bench on @huggingface: the world’s first long-horizon healthcare benchmark for AI agents. 75 real healthcare workflows + 20 …

Model ReleasesDGX agent

Introducing CHI-Bench on @huggingface: the world’s first long-horizon healthcare benchmark for AI agents. 75 real healthcare workflows + 20 apps + 200+ MCP tools + 1,290 skills + process / outcome rew

LLM Agent Based Renewable Energy Forecasting Using Edge and IoT Data A Review of Solar Wind Weather and Grid Aware Decision Support

Local AiDGX agent

arXiv:2605.25141v1 Announce Type: cross Abstract: Reliable forecasting of renewable energy generation is a foundational requirement for grid stability energy trading battery scheduling and carbon awar

Micro-Swarm Locomotion Optimization in Dynamic Flow using Multi-Objective Multi-Agent Reinforcement Learning

SafetyDGX agent

arXiv:2605.25025v1 Announce Type: new Abstract: Coordinating micro-robotic swarms in physiologically realistic, time-dependent fluid environments remains an unsolved challenge for biomedical and envir

Neural Router: Semantic Content Matching for Agentic AI

Model ReleasesDGX agent

arXiv:2605.25701v1 Announce Type: cross Abstract: Large language models (LLMs) can serve as the semantic-matching engine of a content-based publish/subscribe broker for agentic AI across the edge-clou

New on the Engineering Blog: The access and permissions we grant agents should evolve with their capabilities. In our own products, we set t…

Model ReleasesDGX agent

New on the Engineering Blog: The access and permissions we grant agents should evolve with their capabilities. In our own products, we set these parameters through sandboxing, which limits the scope o

Operationalizing Reconstructive Authority: Runtime Construction, Dependency Resolution, and Execution Gating in Autonomous Agent Systems

SafetyDGX agent

arXiv:2605.23935v1 Announce Type: new Abstract: Autonomous agent systems fail not only due to incorrect decisions, but due to executing decisions whose authority no longer holds at runtime. Prior work

Our strategy lead @yeahfortommy was just on stage with CTO of @alibaba_cloud discussing Hermes Agent at the Qwen Conference, check it out: h…

Model ReleasesDGX agent

Our strategy lead @yeahfortommy was just on stage with CTO of @alibaba_cloud discussing Hermes Agent at the Qwen Conference, check it out: https://www.youtube.com/live/r99c3sfgkmc?si=cnuR07ofhO_V69l-&

ProActor: Timing-Aware Reinforcement Learning for Proactive Task Scheduling Agents

SafetyDGX agent

arXiv:2605.24900v1 Announce Type: new Abstract: Proactive task-oriented agents must autonomously anticipate user needs, identify actionable opportunities, and trigger software actions at appropriate m

Q&A with Sundar Pichai about reshaping the information ecosystem with Search changes, putting AI agents in everything, when AI will replace him as CEO, and more (Nilay Patel/The Verge)

IndustryDGX agent

Nilay Patel / The Verge: Q&A with Sundar Pichai about reshaping the information ecosystem with Search changes, putting AI agents in everything, when AI will replace him as CEO, and more — Today, I'm t

Sources: Qualcomm reached a deal with ByteDance to supply millions of ASICs for AI data centers to support AI agents in the Doubao chatbot; QCOM jumps 5%+ (Ian King/Bloomberg)

IndustryDGX agent

Ian King / Bloomberg: Sources: Qualcomm reached a deal with ByteDance to supply millions of ASICs for AI data centers to support AI agents in the Doubao chatbot; QCOM jumps 5%+ — Qualcomm Inc. reached

25 May 2026

DART: Semantic Recoverability for Structured Tool Agents

Local AiDGX agent

arXiv:2605.23311v1 Announce Type: new Abstract: When a structured tool agent fails mid-execution, the runtime faces a dilemma: replaying the entire task is safe but wasteful, while restoring from a lo

DeepSeek V4 Flash IS BACK on Nous Portal for FREE for use in Hermes Agent! Check it out at https://portal.nousresearch.com/manage-subscripti…

Model ReleasesDGX agent

DeepSeek V4 Flash model has been made available again on the Nous Research portal at no cost for use with Hermes Agent applications. Users can access and utilize this model through the Nous portal's s

Just spent a full coding session with Grok Build and honestly? It's right there with Claude. The model is sharp, the agentic flow holds up o…

Model ReleasesDGX agent

Just spent a full coding session with Grok Build and honestly? It's right there with Claude. The model is sharp, the agentic flow holds up on complex tasks, and it has actual personality. Few rough ed

PathNavigate: A Training-Free Pathology Agent with Surprise-Guided Scan and Shared Slide Memory for Whole-Slide Image VQA

Local AiDGX agent

arXiv:2605.23559v1 Announce Type: cross Abstract: Whole-slide image visual question answering (WSI-VQA) frames pathology as an extreme-context search problem: to answer a free-form clinical query, a s

24 May 2026

> I’m now in the LeCun/Marcus camp on LLMs > real programming agents will need world models > not some RLVR shit it’s over

SafetyDGX agent

> I’m now in the LeCun/Marcus camp on LLMs > real programming agents will need world models > not some RLVR shit it’s over The Eternal Sloptember https://geohot.github.io//blog/jekyll/update/2026/05/2

← Previous
1…131132133134135…300
Next →