AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,737 results
11 Jun 2026

SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior

AgentsDGX agent

arXiv:2606.11543v1 Announce Type: new Abstract: Agent Skills augment large language model (LLM) agents with procedural knowledge at inference time, but current benchmarks rarely distinguish what a Ski

10 Jun 2026

Choosing your surface: Antigravity 2.0, Antigravity CLI, Antigravity IDE, or Antigravity SDK

AgentsDGX agent

TL;DR: Antigravity 2.0: A desktop app to orchestrate multiple autonomous agents working in parallel across independent projects. Antigravity CLI: A terminal interface designed for command-line workflo

CollabSkill: Evaluating Human-Agent Collaboration On Real-World Tasks

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2606.09833v1 Announce Type: cross Abstract: AI agents are reshaping the workspace, leading to drastic change of how humans work. Despite the considerable potential of human-agent collaboration b

Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories

AgentsDGX agent

arXiv:2606.11176v1 Announce Type: cross Abstract: Data tells stories that shape society; the data journalist's job is to turn raw information into stories non-experts can trust. A high-quality news fe

Frontier Coding Agents Use Metaprogramming to Adapt to Unfamiliar Programming Languages

Model ReleasesDGX agent

arXiv:2606.10933v1 Announce Type: new Abstract: LLM-based coding agents are usually evaluated in familiar software settings: mainstream languages, common libraries, and public repositories. These benc

HIPIF: Hierarchical Planning and Information Folding for Long-Horizon LLM Agent Learning

AgentsDGX agent

arXiv:2606.10507v1 Announce Type: new Abstract: While Large Language Models (LLMs) have demonstrated strong capabilities as autonomous agents across a wide range of tasks, their performance often degr

9 Jun 2026

A case study of evaluating AI agents on a neuroscience data-to-discovery pipeline

AgentsDGX agent

arXiv:2606.07718v1 Announce Type: new Abstract: Agentic AI tools offer a promising path to automating software development bottlenecks in scientific research pipelines, particularly for stages that ta

Executable World Models for ARC-AGI-3 in the Era of Coding Agents

Model ReleasesDGX agent

arXiv:2605.05138v2 Announce Type: replace Abstract: We evaluate an initial coding-agent system for ARC-AGI-3 in which the agent maintains an executable Python world model, verifies it against previous

MemToolAgent overview with a simple restaurant booking scenario where the agent retrieves similar memories, receives feedback on an invalid time format, and generates a reflection to update its memory

AgentsDGX agent

arXiv:2606.07909v1 Announce Type: new Abstract: Modern large language model (LLM) agents can use external tools to help users solve complex tasks. However, for problems that require learning from long

8 Jun 2026

DuMate-DeepResearch: An Auditable Multi-Agent System with Recursive Search and Rubric-Grounded Reasoning

AgentsDGX agent

arXiv:2606.07299v1 Announce Type: new Abstract: Deep Research (DR) has emerged as a new agentic paradigm to tackle complex, open-ended research tasks, demanding systems that can iteratively frame prob

Measuring Agents in Production

AgentsDGX agent

arXiv:2512.04123v4 Announce Type: replace-cross Abstract: LLM-based agents already operate in production across many industries, yet we lack an understanding of what technical methods make deployments

6 Jun 2026

Beyond Similarity: Trustworthy Memory Search for Personal AI Agents

AgentsDGX agent

arXiv:2606.06054v1 Announce Type: new Abstract: Personal AI agents increasingly rely on long-term memory to provide persistent personalization across sessions. However, existing memory pipelines are l

4 Jun 2026

Designing the hf CLI as an agent-optimized way to work with the Hub

AgentsDGX agent

The Hugging Face CLI (command-line interface) is designed as an agent-optimized tool that enables AI agents and users to interact with the Hugging Face Hub more efficiently. The design prioritizes com

nice post from @Harvey and @LangChain Labs, worth a read. improve agent feedback loop without setting $$ on fire

AgentsDGX agent

nice post from @Harvey and @LangChain Labs, worth a read. improve agent feedback loop without setting $$ on fire Can we design legal agent verifiers that are up to 1,000x cheaper? Verifiers are LLM ju

SMAC-Talk: A Natural Language Extension of the StarCraft Multi-Agent Challenge for Large Language Models

Model ReleasesDGX agent

arXiv:2606.04202v1 Announce Type: new Abstract: As LLMs become more widely deployed, they are increasingly expected to work alongside other AI agents rather than operating in isolation. Effective coor

3 Jun 2026

EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management

AgentsDGX agent

arXiv:2606.03841v1 Announce Type: new Abstract: Recent progress in Large Language Model (LLM) agents has enabled promising advances in automated data science. However, existing approaches remain funda

LAP: An Agent-to-Instrument Protocol for Autonomous Science

SafetyDGX agent

arXiv:2606.03755v1 Announce Type: new Abstract: Autonomous science is moving from demonstration to infrastructure. Large language model agents now plan experiments, and self-driving laboratories execu

The Deliberative Illusion: Diagnosing Factual Attrition and Stance Homogenization in Multi-Agent LLM Deliberation

AgentsDGX agent

arXiv:2606.03032v1 Announce Type: new Abstract: Multi-agent LLM systems often treat consensus as evidence of successful interaction. For deliberative problems, however, reliability depends on whether

2 Jun 2026

Constitutional Black-Box Monitoring for Scheming in LLM Agents

AgentsDGX agent

arXiv:2603.00829v2 Announce Type: replace-cross Abstract: Safe deployment of Large Language Model (LLM) agents in autonomous settings requires reliable oversight mechanisms. A central challenge is det

Coordination Graphs for Constrained Multi-Agent Reinforcement Learning

AgentsDGX agent

arXiv:2606.02337v1 Announce Type: new Abstract: Constrained Multi-agent reinforcement learning (CMARL) faces two intertwined challenges: the joint action space grows exponentially with the number of a

Deliberative Curation: A Protocol for Multi-Agent Knowledge Bases

AgentsDGX agent

arXiv:2606.00007v1 Announce Type: new Abstract: As AI agents transition from isolated tools to collaborative participants in shared knowledge ecosystems, governing collective knowledge curation become

Latent Collaboration in Multi-Agent Systems

AgentsDGX agent

arXiv:2511.20639v3 Announce Type: replace-cross Abstract: Multi-agent systems (MAS) extend large language models (LLMs) from independent single-model reasoning to coordinative system-level intelligenc

Not All Flips Are Conformity: Decomposing Stance Convergence in Multi-Agent LLM Debate

AgentsDGX agent

arXiv:2606.00820v1 Announce Type: new Abstract: Multi-agent debate (MAD) is a promising strategy for improving LLM reasoning, but when agents converge on a shared answer, it is unclear whether that co

On Information Self-Locking in Reinforcement Learning for Active Reasoning of LLM agents

AgentsDGX agent

arXiv:2603.12109v2 Announce Type: replace Abstract: Reinforcement learning (RL) has become a de facto paradigm for building LLM-based agents that act, interact, and reason over extended task horizons.

PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say

Model ReleasesDGX agent

arXiv:2606.00152v1 Announce Type: cross Abstract: LLM-based agents are rapidly advancing, autonomously invoking external tools to complete multi-step tasks for users. However, agents often acquire mor

'Skill issues'': data-centric optimization of lakehouse agents

AgentsDGX agent

arXiv:2606.01185v1 Announce Type: new Abstract: Coding agents are becoming users of data infrastructure, but their success depends not only on model quality: it also depends on the skills and environm

SkillRevise: Improving LLM-Authored Agent Skills via Trace-Conditioned Skill Revision

AgentsDGX agent

arXiv:2606.01139v1 Announce Type: new Abstract: Agent skills are procedural artifacts that enable LLM agents to execute workflows, verify constraints, and recover from failures. Existing self-evolving

1 Jun 2026

A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents

AgentsDGX agent

arXiv:2602.08964v2 Announce Type: replace-cross Abstract: Understanding an agent's goals helps explain and predict its behaviour, yet there is no established methodology for reliably attributing goals

AgentOps: Operationalize agentic AI at scale with Amazon Bedrock AgentCore

AgentsDGX agent

When you build agentic AI solutions, you face unique operational challenges. Agents make unpredictable decisions, costs spiral unexpectedly, and debugging non-deterministic failures seems impossible.

have manually read 1000s traces since joining LangChain! it’s a great way to learn and understand your agent but completely infeasible to do…

AgentsDGX agent

have manually read 1000s traces since joining LangChain! it’s a great way to learn and understand your agent but completely infeasible to do at agent scale 🙃 engine helps us automate that process so h

Scalable Constrained Multi-Agent Reinforcement Learning via State Augmentation and Consensus for Separable Dynamics

SafetyDGX agent

arXiv:2605.30461v1 Announce Type: cross Abstract: We present a distributed approach for constrained Multi-Agent Reinforcement Learning (MARL) that combines state-augmented policy learning with distrib

Why Video Agent models are next — Ethan He, xAI Grok Imagine

AgentsDGX agent

Video agent models represent the next frontier in AI by extending language model capabilities to process, understand, and act on video content in real-time, enabling autonomous agents to perceive and

31 May 2026

It's a big week for Hermes Agent.

AgentsDGX agent

Nous Research announced significant developments or milestones for Hermes Agent, their AI agent framework, indicating multiple new capabilities, updates, or releases during that particular week. The a

29 May 2026

Beyond Consensus: Trace-Level Synthesis in Mixture of Agents

AgentsDGX agent

arXiv:2605.29116v1 Announce Type: new Abstract: When multiple LLM agents solve the same problem, standard practice compresses each agent's reasoning into a majority vote or layered synthesis, treating

V2XCrafter: Learning to Generate Driving Scene Across Agents

SafetyDGX agent

arXiv:2605.29471v1 Announce Type: new Abstract: Collaborative driving systems leverage vehicle-to-everything (V2X) communication for multi-agent collaborative perception to enhance driving safety, yet

WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction

AgentsDGX agent

arXiv:2605.29341v1 Announce Type: cross Abstract: Multimodal large language models are increasingly deployed as long-horizon agents, where memory must do more than recall: it must track an evolving wo

28 May 2026

Agent Explorative Policy Optimization for Multimodal Agentic Reasoning

SafetyDGX agent

arXiv:2605.28774v1 Announce Type: new Abstract: Vision-language models with extended reasoning succeed on complex problems, but many real-world problems require external tools that internal reasoning

AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation

AgentsDGX agent

arXiv:2605.28655v1 Announce Type: new Abstract: Scientific research proceeds through iterative cycles of hypothesis generation, experiment design, execution, and revision. AI agents can automate parts

COOP^2: Defining, Observing, and Repairing Cooperation in LLM Multi-Agent Systems

AgentsDGX agent

arXiv:2603.00349v2 Announce Type: replace Abstract: Many complex tasks require extended effort, diverse capabilities, or coordinated actions beyond what a single agent can provide. However, simply add

Evaluating Deep Agents using LangSmith on AWS

AgentsDGX agent

This post combines learnings from LangChain’s work on evaluating deep agents and Anthropic’s guide to demystifying evals for AI agents into a practical guide. In this post, you will learn how to: 1) a

Governance should lead to more building, not less Agent development is so much more fun when people don't need to ration tokens because one …

AgentsDGX agent

Governance should lead to more building, not less Agent development is so much more fun when people don't need to ration tokens because one runaway agent might blow up the bill Understand spend, cap t

Learn from Weaknesses: Automated Domain Specialization for Small Computer-Use Agents

AgentsDGX agent

arXiv:2605.28775v1 Announce Type: cross Abstract: Computer-use agents (CUAs) have recently made substantial progress, but deploying a separate large expert for each software domain remains expensive.

Mobile-Aptus: Confidence-Driven Proactive and Robust Interaction in MLLM-based Mobile-Using Agents

SafetyDGX agent

arXiv:2605.28629v1 Announce Type: new Abstract: Recent advancements in multimodal large language models (MLLMs) have shown exceptional potential in enabling mobile-using agents to autonomously execute

Out of Sight, Not Out of Mind: Unveiling Latent Attack in Latent-based Multi-Agent Systems

AgentsDGX agent

arXiv:2605.28214v1 Announce Type: cross Abstract: Latent-based multi-agent systems replace parts of explicit inter-agent communication with hidden representations, offering a new direction for efficie

Skill-as-Pseudocode: Refactoring Skill Libraries to Pseudocode for LLM Agents

AgentsDGX agent

arXiv:2605.27955v1 Announce Type: cross Abstract: Markdown skill libraries for LLM agents ship as free-form prose, forcing the agent to re-derive both the input schema and the concrete invocation synt

27 May 2026

A highlight of new deepagents release is delta channels Drastically improves how we store checkpoints for agents

AgentsDGX agent

A highlight of new deepagents release is delta channels Drastically improves how we store checkpoints for agents Deep Agents v0.6 brings Delta channels, reducing checkpoint storage by up to 100x for l

Computer use in Fleet is now in public beta! You can now give your agents in Fleet access to secure virtual machines allowing them to own en…

AgentsDGX agent

Computer use in Fleet is now in public beta! You can now give your agents in Fleet access to secure virtual machines allowing them to own engineering tasks for you. Use these via Fleet agents in the U

Governed Evolution of Agent Runtimes through Executable Operational Cognition

AgentsDGX agent

arXiv:2605.27328v1 Announce Type: cross Abstract: Recent advances in agentic systems increasingly treat code as an executable operational substrate rather than as a disposable output artifact. Prior w

MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation

AgentsDGX agent

arXiv:2605.27366v1 Announce Type: new Abstract: Large language model (LLM) agents rely on reusable skills to solve complex tasks. However, existing skill creation approaches treat skills as isolated a

Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History

Model ReleasesDGX agent

arXiv:2602.17003v2 Announce Type: replace-cross Abstract: Large language models have advanced web agents, yet current agents lack personalization capabilities. Since users rarely specify every detail

Securing Multi-Agent Systems Against Corruptions via Node Contribution Backpropagation

SafetyDGX agent

arXiv:2510.19420v2 Announce Type: replace-cross Abstract: Multi-Agent Systems (MAS) have become a prevalent paradigm for Large Language Model (LLM) applications. However, the complex multi-agent desig

26 May 2026

DemoEvolve: Overcoming Sparse Feedback in Agentic Harness Evolution with Demonstrations

AgentsDGX agent

arXiv:2605.24539v1 Announce Type: new Abstract: Agent harness evolution improves frozen language-model agents by modifying the executable structures around them. We study this paradigm as a form of sa

// Language Models Need Sleep // Let your agents 'sleep', folks. On a serious note, this is a fascinating paper on getting the most from lon…

TutorialsDGX agent

// Language Models Need Sleep // Let your agents 'sleep', folks. On a serious note, this is a fascinating paper on getting the most from long-horizon agents. Here is the problem with agents today: Att

SPARK: Search Personalization via Agent-Driven Retrieval and Knowledge-sharing

AgentsDGX agent

arXiv:2512.24008v3 Announce Type: replace Abstract: Personalized search demands the ability to model users' evolving, multi-dimensional information needs; a challenge for systems constrained by static

25 May 2026

MemAudit: Post-hoc Auditing of Poisoned Agent Memory via Causal Attribution and Structural Anomaly Detection

AgentsDGX agent

arXiv:2605.23723v1 Announce Type: new Abstract: Large language model agents increasingly rely on persistent memory to store past interactions, retrieve relevant demonstrations, and improve long-horizo

When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems

AgentsDGX agent

arXiv:2605.23414v1 Announce Type: new Abstract: LLM-based multi-agent systems can fail even when planned actions are executed correctly because agents may misjudge their knowledge when evaluating plan

23 May 2026

Every 'self-evolving agent' paper this year has mutated text: prompts, skill files, workflow graphs, memory schemas. MOSS from USTC & HKUST …

AgentsDGX agent

Every 'self-evolving agent' paper this year has mutated text: prompts, skill files, workflow graphs, memory schemas. MOSS from USTC & HKUST argues this is the wrong layer. The thing that actually brea

21 May 2026

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs

AgentsDGX agent

arXiv:2605.20315v1 Announce Type: new Abstract: LLM agents have recently emerged as a powerful paradigm for solving complex tasks through planning, tool use, memory retrieval, and multi-step interacti

20 May 2026

CANTANTE: Optimizing Agentic Systems via Contrastive Credit Attribution [R]

AgentsDGX agent

CANTANTE addresses the challenge of optimizing LLM-based multi-agent systems where system-level performance scores are available but individual agent parameters cannot be directly optimized. The frame

Railway: The Agent-Native Cloud — Jake Cooper

AgentsDGX agent

Railway is a cloud platform that provides agent-native infrastructure, enabling AI agents to be deployed and executed natively within the cloud environment. The discussion with Jake Cooper likely cove

← Previous
1…3536373839…296
Next →