AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,964 results
24 May 2026

Q: How are job postings for software engineers rising rapidly despite AI agents automating coding? A: Because there’s far more code to manag…

IndustryDGX agent

Q: How are job postings for software engineers rising rapidly despite AI agents automating coding? A: Because there’s far more code to manage than ever before. We’re already seeing a 14x YoY increase

23 May 2026

Dynamic Mixture of Latent Memories for Self-Evolving Agents

ResearchDGX agent

arXiv:2605.21951v1 Announce Type: new Abstract: Achieving self-evolution in intelligent agents requires the continual accumulation of new knowledge across changing task sequences without forgetting pr

MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2602.13372v2 Announce Type: replace-cross Abstract: Evaluating moral alignment in agents navigating conflicting, hierarchically structured human norms is a critical challenge at the intersection

NEW: Google claimed this week that a team of agents had built an entire operating system, based on a single prompt, costing only about $900 …

SafetyDGX agent

NEW: Google claimed this week that a team of agents had built an entire operating system, based on a single prompt, costing only about $900 in tokens. We fact check this claim and analyze what it mean

Q&A with Sundar Pichai on the future of Google Search, Google's place in the AI race, public skepticism toward AI, AI agents, AI safety, TPUs, and more (New York Times)

SafetyDGX agent

New York Times: Q&A with Sundar Pichai on the future of Google Search, Google's place in the AI race, public skepticism toward AI, AI agents, AI safety, TPUs, and more — After a busy Google I/O, the c

22 May 2026

Big fan of teaching more people the basics of using Claude Code in an accessible way. So much of the world has not yet used agents. There's …

Model ReleasesDGX agent

Big fan of teaching more people the basics of using Claude Code in an accessible way. So much of the world has not yet used agents. There's a lot of opportunity to level the playing field and expand a

Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines

Model ReleasesDGX agent

arXiv:2605.20630v1 Announce Type: new Abstract: Industrial asset operations workflows are latency-sensitive because a single user query may require coordination over sensor data, work orders, failure

Personality Engineering with AI Agents: A New Methodology for Negotiation Research

TutorialsDGX agent

arXiv:2605.20554v1 Announce Type: new Abstract: According to canonical negotiation theory, people's success in a negotiation depends on how well they balance competing demands--empathizing and asserti

RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems

Model ReleasesDGX agent

arXiv:2510.13910v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) mitigates key limitations of Large Language Models (LLMs)-such as factual errors, outdated knowledge, and hallu

Retaining Suboptimal Actions to Follow Shifting Optima in Multi-Agent Reinforcement Learning

SafetyDGX agent

arXiv:2602.17062v2 Announce Type: replace Abstract: Value decomposition is a core approach for cooperative multi-agent reinforcement learning (MARL). However, existing methods still rely on a single o

With the Cursor SDK, you can build your own agents with Composer 2.5. It's now available in Python and TypeScript. This long weekend, Compos…

ToolsDGX agent

With the Cursor SDK, you can build your own agents with Composer 2.5. It's now available in Python and TypeScript. This long weekend, Composer usage is 90% off in the SDK. We're excited to see what yo

21 May 2026

Everyone on earth will get really good chatbots for free, so thats good for democratization of AI, but really good agents that can do comple…

ApplicationsDGX agent

Everyone on earth will get really good chatbots for free, so thats good for democratization of AI, but really good agents that can do complex work burn thousands of times more tokens, and will be rese

I released the first alpha of Datasette Agent - a conversational AI assistant for Datasette that can answer questions about data in SQLite d…

Model ReleasesDGX agent

I released the first alpha of Datasette Agent - a conversational AI assistant for Datasette that can answer questions about data in SQLite databases, and can be extended with plugins to add extra tool

In the next version of Claude Code: run /usage to see a breakdown of which Skills, Agents, MCPs, and Plugins are using your tokens CLI today…

Model ReleasesDGX agent

The next version of Claude Code will introduce a `/usage` command that provides a detailed breakdown of token consumption across different components including Skills, Agents, MCPs (Model Context Prot

Multi-Agent Reinforcement Learning for Safe Autonomous Driving Under Pedestrian Behavioral Uncertainty

SafetyDGX agent

arXiv:2605.20255v1 Announce Type: new Abstract: Simulation-based testing of self-driving cars (SDCs) typically relies on scripted or simplified pedestrian models that do not capture the heterogeneity

STEAM: A Training-Free Congestion-Aware Enhancement Framework for Decentralized Multi-Agent Path Finding

SafetyDGX agent

arXiv:2605.20929v1 Announce Type: new Abstract: We propose STEAM (Spatial, Temporal, and Emergent congestion Awareness for MAPF), a training-free test-time enhancement framework for learning-based dec

The Yes-Man Syndrome: Benchmarking Abstention in Embodied Robotic Agents

Model ReleasesDGX agent

arXiv:2605.20544v1 Announce Type: cross Abstract: Vision-language models (VLMs) are used as high-level planners for embodied agents, translating natural language instructions and visual observations i

TLMs: Tiny LLMs and Agents on Edge Devices with @cormacb https://www.youtube.com/watch?v=-TiET_K-E_g Function Gemma ships at 270 million par…

Model ReleasesDGX agent

TLMs: Tiny LLMs and Agents on Edge Devices with @cormacb https://www.youtube.com/watch?v=-TiET_K-E_g Function Gemma ships at 270 million parameters and runs nearly 2,000 tokens per second prefill on a

20 May 2026

4 big upgrades to Hermes Agents speed today `hermes update` to get moving faster now

ResearchDGX agent

Nous Research announced four major performance upgrades to their Hermes Agents framework, aimed at improving speed and efficiency. The update, referred to as the `hermes update`, enables faster execut

ContextFlow: Hierarchical Task-State Alignment for Long-Horizon Embodied Agents

SafetyDGX agent

arXiv:2605.19314v1 Announce Type: cross Abstract: Long-horizon embodied agents increasingly delegate navigation, search, approach, and manipulation to specialist executors. As these executors become s

DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows

Model ReleasesDGX agent

arXiv:2605.19099v1 Announce Type: new Abstract: We introduce DecisionBench, a benchmark substrate for emergent delegation in long-horizon agentic workflows. The substrate fixes a task suite (GAIA, tau

Evaluating Memory Condensation Strategies for Coding Agents in Data-Driven Scientific Discovery

TutorialsDGX agent

arXiv:2605.18854v1 Announce Type: new Abstract: Coding agents accumulate extensive context during long-running tasks, yet fixed context windows force practitioners to choose between truncation and tas

I'm switching to Hermes.... I've been using it for a month.....and I'm sold...moving all of my @openclaw agents to Hermes (@NousResearch) Wh…

ResearchDGX agent

I'm switching to Hermes.... I've been using it for a month.....and I'm sold...moving all of my @openclaw agents to Hermes (@NousResearch) Why? -----> https://youtu.be/QQEgIo4Juxg Thank you to @Hosting

Memory-Augmented Reinforcement Learning Agent for CAD Generation

SafetyDGX agent

arXiv:2605.19748v1 Announce Type: new Abstract: Automatic generation of computer-aided design (CAD) models is a core technology for enabling intelligence in advanced manufacturing. Existing generation

MiniMax Speech 2.8 Turbo is built for voice agents that need natural delivery, not just clean audio. → Sound Tags for laughter, breathing, s…

ToolsDGX agent

MiniMax Speech 2.8 Turbo is built for voice agents that need natural delivery, not just clean audio. → Sound Tags for laughter, breathing, sighs, gasps, and other vocal cues → 60% prosody improvement

Phase-Aware Mixture of Experts for Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2602.17038v3 Announce Type: replace Abstract: Reinforcement learning (RL) has equipped LLM agents with a strong ability to solve complex tasks. However, existing RL methods normally use a single

Rethinking How to Remember: Beyond Atomic Facts in Lifelong LLM Agent Memory

Model ReleasesDGX agent

arXiv:2605.19952v1 Announce Type: new Abstract: To enable reliable long-term interaction, LLM agents require a memory system that can faithfully store, efficiently retrieve, and deeply reason over acc

SimGym: A Framework for A/B Test Simulation in E-Commerce with Traffic-Grounded VLM Agents

SafetyDGX agent

arXiv:2605.19219v1 Announce Type: new Abstract: A/B testing remains the gold standard for evaluating modifications to e-commerce storefronts, yet it diverts traffic, requires weeks to reach statistica

To Call or Not to Call: Diagnosing Intrinsic Over-Calling Bias in LLM Agents

Model ReleasesDGX agent

arXiv:2605.18882v1 Announce Type: cross Abstract: LLM agents exhibit a consistent tendency to over-call, invoking tools even in situations where none is needed. On the When2Call benchmark, six models

19 May 2026

Albert Einstein + ElevenLabs. AI agents can make education more accessible - a teacher for every student in every field. A classroom size of…

ApplicationsDGX agent

Albert Einstein + ElevenLabs. AI agents can make education more accessible - a teacher for every student in every field. A classroom size of one learning from icons who shaped the world Today with his

BLAgent: Agentic RAG for File-Level Bug Localization

Local AiDGX agent

arXiv:2605.17965v1 Announce Type: cross Abstract: Bug localization remains a key bottleneck in downstream software maintenance tasks, including root cause analysis, triage, and automated program repai

Causal Intervention-Based Memory Selection for Long-Horizon LLM Agents

Model ReleasesDGX agent

arXiv:2605.17641v1 Announce Type: new Abstract: Long-horizon LLM agents rely on persistent memory to support interactions across sessions, yet existing memory systems often retrieve context using sema

Computer use turns Claude into an agent that can operate real UIs. New blog post on making it reliable in production: getting click accuracy…

Model ReleasesDGX agent

Computer use turns Claude into an agent that can operate real UIs. New blog post on making it reliable in production: getting click accuracy right, choosing thinking effort levels, keeping long sessio

Enhancing Cloud Network Resilience via a Robust LLM-Empowered Multi-Agent Reinforcement Learning Framework

Local AiDGX agent

arXiv:2601.07122v2 Announce Type: replace-cross Abstract: While virtualization and resource pooling empower cloud networks with structural flexibility and elastic scalability, they inevitably expand t

Equilibrium Selection in Multi-Agent Policy Gradients via Opponent-Aware Basin Entry

SafetyDGX agent

arXiv:2605.18078v1 Announce Type: new Abstract: Multi-agent policy-gradient methods have been shown to converge locally near stable Nash equilibria. Local convergence, however, does not determine whic

extsc{PrivScope}: Task-scoped Disclosure Control for Hybrid Agentic Systems

Model ReleasesDGX agent

arXiv:2605.16630v1 Announce Type: cross Abstract: Hybrid local--cloud agents enrich user requests with context from persistent working state before delegating capability-intensive subtasks to a cloud

Generation Navigator: A State-Aware Agentic Framework for Image Generation

SafetyDGX agent

arXiv:2605.17969v1 Announce Type: new Abstract: Despite rapid advances in text-to-image generation, faithfully realizing user intent remains challenging, often requiring manual multi-turn trial and er

Introducing Claude Managed Agents with Modal Sandboxes

Model ReleasesDGX agent

This article announces a collaboration between Anthropic's Claude and Modal that enables Claude to operate as managed agents within Modal's sandboxed environments. The integration allows Claude to saf

LaunchDarkly launches runtime control layer for the agentic AI era

Model ReleasesDGX agent

LaunchDarkly, a feature control platform that helps developers and software engineers launch and manage products, today announced the launch of AgentControl, a new solution providing real-time managem

MAVEN A Multi-Agent Framework for Multicultural Text-to-Video Generation

Model ReleasesDGX agent

arXiv:2605.16716v1 Announce Type: cross Abstract: Text-to-video (T2V) generation has rapidly progressed in visual fidelity, yet its ability to faithfully represent multiple cultures within a single pr

Mitigating Conversational Inertia in Multi-Turn Agents

SafetyDGX agent

arXiv:2602.03664v3 Announce Type: replace Abstract: Large language models excel as few-shot learners when provided with appropriate demonstrations, yet this strength becomes problematic in multiturn a

Robo-Cortex: A Self-Evolving Embodied Agent via Dual-Grain Cognitive Memory and Autonomous Knowledge Induction

Local AiDGX agent

arXiv:2605.18729v1 Announce Type: cross Abstract: The ability to navigate and interact with complex environments is central to real-world embodied agents, yet navigation in unseen environments remains

Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents

Model ReleasesDGX agent

arXiv:2605.16986v1 Announce Type: cross Abstract: LLM agents benefit from reusable skills, yet test-time tasks often require guidance more specific than a static skill library can provide. We propose

SPIKE: An Adaptive Dual Controller Framework for Cost-Efficient Long-Horizon Game Agents

ResearchDGX agent

arXiv:2605.18636v1 Announce Type: new Abstract: Long-horizon multimodal agents in open-world games must stay goal-directed across many low-level interactions under tight token and latency budgets. Exi

TusoAI: Agentic Optimization for Scientific Methods

Model ReleasesDGX agent

arXiv:2509.23986v2 Announce Type: replace Abstract: Scientific discovery is often slowed by the manual development of computational tools needed to analyze complex experimental data. Building such too

WEBSERV: A Full-Stack and RL-Ready Web Environment for Training Web Agents at Scale

Model ReleasesDGX agent

arXiv:2510.16252v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) for web agents demands environments that are both effective for evaluation and efficient enough for large-scale on

18 May 2026

A mental model for working with coding agents is that they're blind squirrels running into a maze and bumping into walls. You must place the…

ResearchDGX agent

A mental model for working with coding agents is that they're blind squirrels running into a maze and bumping into walls. You must place the walls (verifiable constraints) strategically so that they e

Companies running bug bounty programs are tightening background checks and building AI agents to triage a flood of low-quality reports generated by AI (Jamie John/Financial Times)

IndustryDGX agent

Jamie John / Financial Times: Companies running bug bounty programs are tightening background checks and building AI agents to triage a flood of low-quality reports generated by AI — ‘Bug bounty’ prog

TopoEvo: A Topology-Aware Self-Evolving Multi-Agent Framework for Root Cause Analysis in Microservices

SafetyDGX agent

arXiv:2605.15611v1 Announce Type: new Abstract: Root cause analysis (RCA) in microservices is challenging due to (i) noisy and heterogeneous multimodal observability (metrics, logs, traces), (ii) casc

17 May 2026

We built TERMS-Bench, a three-tier benchmark for LLM agents in real-world economic negotiation. No LLM-as-judge, no outcome rubrics: the env…

Model ReleasesDGX agent

We built TERMS-Bench, a three-tier benchmark for LLM agents in real-world economic negotiation. No LLM-as-judge, no outcome rubrics: the environment itself is the verifier. 🏆Among frontier models, @An

16 May 2026

LLM-powered AI agents are gonna be great! You should totally trust them!

SafetyDGX agent

Gary Marcus expresses optimism about the potential of LLM-powered AI agents in a post on X (formerly Twitter). The post advocates for confidence in these AI systems, though without access to the full

15 May 2026

AgenticEval: Toward Agentic and Self-Evolving Safety Evaluation of Large Language Models

Model ReleasesDGX agent

arXiv:2509.26100v2 Announce Type: replace Abstract: The rapid integration of Large Language Models (LLMs) into high-stakes domains necessitates reliable safety and compliance evaluation. However, exis

EvolveMem:Self-Evolving Memory Architecture via AutoResearch for LLM Agents

Model ReleasesDGX agent

arXiv:2605.13941v1 Announce Type: cross Abstract: Long-term memory is essential for LLM agents that operate across multiple sessions, yet existing memory systems treat retrieval infrastructure as fixe

From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents

Model ReleasesDGX agent

arXiv:2605.15104v1 Announce Type: new Abstract: Voice agents increasingly require reliable tool use from speech, whereas prominent tool-calling benchmarks remain text-based. We study whether verified

Latency-Quality Routing for Functionally Equivalent Tools in LLM Agents

Model ReleasesDGX agent

arXiv:2605.14241v1 Announce Type: new Abstract: Tool-augmented LLM agents increasingly access the same tool type through multiple functionally equivalent providers, such as web-search APIs, retrievers

MemReranker: Reasoning-Aware Reranking for Agent Memory Retrieval

Model ReleasesDGX agent

arXiv:2605.06132v2 Announce Type: replace Abstract: In agent memory systems, the reranking model serves as the critical bridge connecting user queries with long-term memory. Most systems adopt the 're

Probabilistic Verification of Recurrent Neural Networks for Single and Multi-Agent Reinforcement Learning

SafetyDGX agent

arXiv:2605.14758v1 Announce Type: new Abstract: History-dependent policies induced by recurrent neural networks (RNNs) rely on latent hidden state dynamics, making verification in partially observable

Reinforcement Learning for Tool-Calling Agents in Fast Healthcare Interoperability Resources (FHIR)

Model ReleasesDGX agent

arXiv:2605.14126v1 Announce Type: cross Abstract: Fast Healthcare Interoperability Resources (FHIR) is the dominant standard for interoperable exchange of healthcare data. In FHIR, electronic health r

Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy

SafetyDGX agent

arXiv:2605.14558v1 Announce Type: cross Abstract: Agentic reinforcement learning trains large language models using multi-turn trajectories that interleave long reasoning traces with short environment

Towards In-Depth Root Cause Localization for Microservices with Multi-Agent Recursion-of-Thought

Local AiDGX agent

arXiv:2605.14866v1 Announce Type: cross Abstract: As modern microservice systems grow increasingly complex due to dynamic interactions and evolving runtime environments, they experience failures with

← Previous
1…132133134135136…300
Next →