AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,762 results
5 Jul 2026

Active Graphの「The Log Is the Agent」で何が変わる? AIエージェントを“会話”ではなく“ログ”で動かす発想です。 監査・差し戻し・再実行が重いチームほど、この設計が効きます。

AgentsDGX agent

Active Graph proposes a paradigm shift in AI agent architecture by using logs as the primary mechanism for agent operations rather than conversation-based interactions. This design approach is particu

4 Jul 2026

Sakana AI is heading to #ICML2026 in Seoul (July 6–11)! 🐟🇰🇷 Our team will present 11 papers spanning multi-agent coordination, sparse and…

AgentsDGX agent

Sakana AI is heading to #ICML2026 in Seoul (July 6–11)! 🐟🇰🇷 Our team will present 11 papers spanning multi-agent coordination, sparse and efficient LLMs, test-time scaling, long-term memory, and agent

3 Jul 2026

GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2606.22737v2 Announce Type: replace Abstract: Before letting an agent operate over real context, can you prove it used the right evidence? GroundEval turns that question into a deterministic tes

SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use

AgentsDGX agent

arXiv:2607.01874v1 Announce Type: new Abstract: Skills are becoming a reusable operational layer for LLM agents, encoding SOPs, domain rules, tool workflows, scripts, and validation routines. In reali

Vercel's Andrew Qu on why agents are a new kind of software

AgentsDGX agent

Andrew Qu from Vercel discusses how AI agents represent a fundamentally different category of software compared to traditional applications, exploring their unique characteristics and implications for

2 Jul 2026

Multi-Agent Teams Hold Experts Back

AgentsDGX agent

Multi-agent LLM systems are increasingly deployed as autonomous collaborators, where agents interact freely rather than execute fixed, pre-specified workflows. In such settings, effective coordination

SlowBA: An efficiency backdoor attack towards VLM-based GUI agents

AgentsDGX agent

arXiv:2603.08316v3 Announce Type: replace-cross Abstract: Modern vision-language-model (VLM) based graphical user interface (GUI) agents are expected not only to execute actions accurately but also to

1 Jul 2026

Beyond expert users: agents should help users construct preferences, not just elicit them

Model ReleasesDGX agent

arXiv:2606.30863v1 Announce Type: new Abstract: Agents typically assume an expert user -- one with well-formed preferences about what they want -- and default to clarifying questions whenever the task

the log is the agent!

AgentsDGX agent

This post explores the concept that an agent's log or execution history serves as the primary mechanism for decision-making and learning, suggesting that the sequential record of actions and outcomes

30 Jun 2026

Agent-Computer Observation Interfaces Enable Dynamic Computer Use

Model ReleasesDGX agent

arXiv:2606.29472v1 Announce Type: new Abstract: SWE-agent established the action interface as an underexplored design axis for software-engineering agents; we make the analogous case for the observati

Agentic AI for ISAC: Analysis, Framework, and Case Study

AgentsDGX agent

arXiv:2512.15044v2 Announce Type: replace Abstract: Integrated sensing and communication (ISAC) has emerged as a key development direction in the sixth-generation (6G) era, which provides essential su

Agentic AI Health Assistants for Proactive Patient Care and Scalable Risk Management

AgentsDGX agent

This article likely explores how agentic AI systems can proactively monitor patient health data and manage clinical risks at scale, improving healthcare delivery through autonomous AI agents. It proba

agents that can write code can solve problems more reliably but you need to make sure you execute that untrusted code in a safe environment …

AgentsDGX agent

Code-writing AI agents can solve complex problems more effectively than agents using other approaches, but deploying them requires executing potentially untrusted generated code within sandboxed or is

Capability Gates Are Not Authorization: Confused-Deputy Failures in LLM Agent Frameworks

AgentsDGX agent

arXiv:2606.28679v1 Announce Type: cross Abstract: Tool-using LLM agents increasingly read untrusted content while holding side-effecting tools such as payments, email, CRM, and infrastructure APIs, ye

Demonstration-Free Robotic Control via LLM Agents

Model ReleasesDGX agent

arXiv:2601.20334v2 Announce Type: replace-cross Abstract: Robotic manipulation has increasingly adopted vision-language-action (VLA) models, which achieve strong performance but typically require task

Evaluate agents with Harbor + LangChain ‼️

AgentsDGX agent

This post likely discusses how to evaluate AI agents built with LangChain using Harbor, a tool for testing and monitoring. It covers practical approaches for assessing agent performance, reliability,

Experience Graphs: The Data Foundation for Self-Improving Agents

AgentsDGX agent

arXiv:2606.29823v1 Announce Type: cross Abstract: The database community has repeatedly advanced the state of the art by recognizing that new workloads demand new system architectures. We argue that l

Multi-Agent Routing as Set-Valued Prediction: A WildChat Benchmark and Cost-Aware Evaluation

Model ReleasesDGX agent

arXiv:2606.28925v1 Announce Type: cross Abstract: Tool and agent routing from natural-language prompts is naturally a set-valued prediction problem: a single query may require multiple agents, while o

SEATauBench: Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages

Model ReleasesDGX agent

arXiv:2606.28715v1 Announce Type: cross Abstract: While AI development and evaluation for Southeast Asia (SEA) has grown rapidly, agent capabilities in regional languages are still poorly understood d

TraceLab: Characterizing Coding Agent Workloads for LLM Serving

Model ReleasesDGX agent

arXiv:2606.30560v1 Announce Type: cross Abstract: Coding agents are rapidly becoming a major application of agentic LLMs, but serving them efficiently remains challenging. Progress on this challenge r

Vercel Agent has updated pricing

AgentsDGX agent

Vercel Agent pricing has been updated, likely reflecting changes to the cost structure for using Vercel's AI agent capabilities. The specific details of the new pricing tiers, feature inclusions, and

What Capable Agents Must Know: Selection Theorems for Robust Decision-Making under Uncertainty

AgentsDGX agent

arXiv:2603.02491v3 Announce Type: replace-cross Abstract: As artificial agents become increasingly capable, what internal structure is necessary for an agent to act competently under uncertainty? Clas

29 Jun 2026

Baz releases Baz Planner, which uses four specialized AI agents to analyze code at the planning stage, and extends its seed funding by 9M to 17M (Mike Wheatley/SiliconANGLE)

AgentsDGX agent

Mike Wheatley / SiliconANGLE: Baz releases Baz Planner, which uses four specialized AI agents to analyze code at the planning stage, and extends its seed funding by 9M to 17M — Agentic coding startup

Contagion Networks: Evaluator Preference Propagation in Multi-Agent LLM Systems

Model ReleasesDGX agent

arXiv:2606.20493v2 Announce Type: replace-cross Abstract: When large language models serve as evaluators in multi-agent systems, their strategy preferences -- whether induced by explicit prompts or by

Debugging production agents with Amazon Bedrock AgentCore Observability

AgentsDGX agent

In this post, you learn how to debug production agent failures using built-in observability capabilities. We walk through common failure patterns, show how to analyze agent behavior with traces and me

Hierarchical Control in Multi-Agent Games: LLM-based Planning and RL Execution

AgentsDGX agent

arXiv:2606.20014v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has achieved strong performance in sequential decision-making, yet scaling to complex multi-agent environments rem

i've been using cursor mobile on the go for the last weeks, and having access to all cloud agents from everywhere is really nice go on a wal…

AgentsDGX agent

i've been using cursor mobile on the go for the last weeks, and having access to all cloud agents from everywhere is really nice go on a walk, get an idea, dictate it in the app come back from walk to

LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior

Model ReleasesDGX agent

arXiv:2606.28182v1 Announce Type: cross Abstract: Embodied agents operating in decentralized and partially observable environments have attracted growing attention in recent years. However, existing l

NEW paper from Google (bookmark it) It's on advancing automated scientific review. Just pay attention to the focus on agentic verification w…

AgentsDGX agent

NEW paper from Google (bookmark it) It's on advancing automated scientific review. Just pay attention to the focus on agentic verification which is something I've been writing about recently. AI is ac

Row-Bot runs a LangGraph agent at its core. And supports 2 voice pipelines: - Local STT using whisper + Local TTS using kokoro - A real-time…

AgentsDGX agent

Row-Bot runs a LangGraph agent at its core. And supports 2 voice pipelines: - Local STT using whisper + Local TTS using kokoro - A real-time voice pipeline using GPT realtime 2. Both support full agen

26 Jun 2026

💪deep agents harness

AgentsDGX agent

💪deep agents harness The AI SDK Harness API now supports @OpenCode and @LangChain Deep Agents through a single unified interface. Use 𝙷𝚊𝚛𝚗𝚎𝚜𝚜𝙰𝚐𝚎𝚗𝚝 to run any supported runtime without changing your ap

The Red Queen Godel Machine: Co-Evolving Agents and Their Evaluators

Model ReleasesDGX agent

arXiv:2606.26294v1 Announce Type: cross Abstract: Self-improving agents are state-of-the-art (SOTA) on agentic coding benchmarks and have recently been extended to general domains. However, their sear

25 Jun 2026

Diagnosing and Mitigating Compounding Failures in Agentic Persuasion via Taxonomic Strategy Retrieval

AgentsDGX agent

arXiv:2606.24976v1 Announce Type: cross Abstract: Foundation-model agents in multi-step, open-ended environments frequently suffer from compounding errors, where early mistakes contaminate long-horizo

excited to be speaking at @aiDotEngineer World Fair next week on Improving Agents, Continual Learning, and why we think a large part of it i…

AgentsDGX agent

excited to be speaking at @aiDotEngineer World Fair next week on Improving Agents, Continual Learning, and why we think a large part of it is...Data Mining! Trace Mining is how we understand agent beh

24 Jun 2026

Debate2Create: Robot Co-design via Multi-Agent LLM Debate

AgentsDGX agent

arXiv:2510.25850v3 Announce Type: replace-cross Abstract: We introduce Debate2Create (D2C), a multi-agent LLM framework that formulates robot co-design as structured, iterative debate grounded in phys

Learn *anything* with our new /learn agent skill.

AgentsDGX agent

Learn *anything* with our new /learn agent skill. Obsessed with our new /learn skill. It's my favorite way of learning and researching topics. The agent creates a learning plan and a learning hub (art

📣📣 Meet Qwen-AgentWorld — a native language world model that simulates 7 agent environments (MCP, Search, Terminal, SWE, Web, OS, Android)…

Model ReleasesDGX agent

📣📣 Meet Qwen-AgentWorld — a native language world model that simulates 7 agent environments (MCP, Search, Terminal, SWE, Web, OS, Android) within a single model. Environment modeling is the training o

The next scaling law is multi-agent swarms. Mixture of models.

AgentsDGX agent

The next scaling law is multi-agent swarms. Mixture of models. Introducing Sakana Fugu: A full multi-agent orchestration system accessible via a single model API. Our ‘Fugu Ultra’ model matches the pe

23 Jun 2026

ActiveGraph: 1 month in: 📄Paper #1: The Log is the Agent 🧠3 LongMemEval Experiments 🔄 Paper #2: Regimes, self-improvement loop 🎓 http://…

AgentsDGX agent

ActiveGraph: 1 month in: 📄Paper #1: The Log is the Agent 🧠3 LongMemEval Experiments 🔄 Paper #2: Regimes, self-improvement loop 🎓 http://learn.activegraph.ai 🗂️ 2 reference agents (code, research) 💾 co

GLM-5.2 is available in Perplexity's Agent API. Just tested it, and it's powerful when paired with the Search SDK inside a sandbox. - Spin u…

AgentsDGX agent

GLM-5.2 is available in Perplexity's Agent API. Just tested it, and it's powerful when paired with the Search SDK inside a sandbox. - Spin up a sandbox environment - Call the web search tool (people s

I'm digging the eve agentic framework from Vercel. I like that everything is files, from the tools to the skills to the evals. More importan…

AgentsDGX agent

I'm digging the eve agentic framework from Vercel. I like that everything is files, from the tools to the skills to the evals. More importantly, it's gets you building with agents fast. Very promising

Memory Contagion: Cross-Temporal Propagation of Evaluator Bias via Agent Memory

SafetyDGX agent

arXiv:2606.23195v1 Announce Type: new Abstract: Large Language Model (LLM) agents increasingly rely on memory systems to maintain long-term coherence. Recent work shows that agent memories degrade dur

Steer, Don't Solve: Training Small Critic Models for Large Code Agents

Model ReleasesDGX agent

arXiv:2606.21811v1 Announce Type: cross Abstract: End-to-end code agent training is resource-intensive and plateaus on the strategy-level reasoning needed to resolve code issues, since jointly optimiz

20 Jun 2026

Hermes Agent has a new Blank Slate setup mode. The default Quick/Full setup modes work great for most, but if you would rather build your ag…

AgentsDGX agent

Hermes Agent has a new Blank Slate setup mode. The default Quick/Full setup modes work great for most, but if you would rather build your agent from the ground up you can now start with just a provide

11 Jun 2026

AI Coding Agents Can Reproduce Social Science Findings

Model ReleasesDGX agent

arXiv:2606.11447v1 Announce Type: new Abstract: Recent anecdotal evidence suggests that AI coding agents can reproduce published findings when provided with original data and code; yet systematic eval

Knowing When to Ask: Self-Gated Clarification for Hierarchical Language Agents

AgentsDGX agent

arXiv:2606.11349v1 Announce Type: new Abstract: In hierarchical reasoning, failures often originate at intermediate decision points where the agent commits to a wrong branch without recognizing that i

Organize then Retrieve: Hierarchical Memory Navigation for Efficient Agents

AgentsDGX agent

arXiv:2606.11680v1 Announce Type: new Abstract: Large language model (LLM) agents struggle with long-horizon tasks due to their inherent statelessness, requiring all task-relevant information to be en

10 Jun 2026

The Confident Liar: Diagnosing Multi-Agent Debate with Log-Probabilities and LLM-as-Judge

AgentsDGX agent

arXiv:2606.10296v1 Announce Type: cross Abstract: Multi-agent debate systems are typically evaluated only on whether the final answer is correct, overlooking the quality of the intermediate reasoning

VISTA: A Versatile Interactive User Simulation Toolkit for Agent Evaluation

AgentsDGX agent

arXiv:2606.11079v1 Announce Type: new Abstract: Evaluation remains a critical bottleneck for interactive agent development. Existing evaluation methods often rely on static benchmarks, which fail to c

9 Jun 2026

How an Agent Built a 3D Paris Gallery by Chaining Two Hugging Face Spaces

AgentsDGX agent

This article describes how an AI agent was used to create a 3D virtual gallery of Paris by chaining together two Hugging Face Spaces applications. It demonstrates a practical example of using agents t

If people only knew how much OpenMed runs on HF Stack like Buckets, Datasets, and Spaces, From datasets, to agent traces, to medical intelli…

AgentsDGX agent

If people only knew how much OpenMed runs on HF Stack like Buckets, Datasets, and Spaces, From datasets, to agent traces, to medical intelligent MCPs, to binary builds for OpenMed Agent, @huggingface

MetaEvo: A Meta-Optimization Framework for Experience-Driven Agent Evolution

AgentsDGX agent

arXiv:2606.07603v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit strong reasoning capabilities, yet most LLM-based agents are statically deployed and unable to improve through ta

PACE: Anytime-Valid Acceptance Tests for Self-Evolving Agents

AgentsDGX agent

arXiv:2606.08106v1 Announce Type: new Abstract: Self-evolving agents improve by repeatedly proposing changes to their own prompts, skills, or workflows and keeping those that score higher on a small h

SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?

Model ReleasesDGX agent

arXiv:2606.07682v1 Announce Type: cross Abstract: AI agents are increasingly expected to complete long-horizon workflows that require sustained progress over hours, millions of tokens, and complex env

The Token Not Taken: Sampling, State, and the Variability of AI Agent Outputs

AgentsDGX agent

arXiv:2606.08998v1 Announce Type: new Abstract: Agentic AI systems can behave differently across runs: the same request may produce a different plan, a different tool call, a different code edit, or a

8 Jun 2026

AdMem: Advanced Memory for Task-solving Agents

AgentsDGX agent

arXiv:2606.06787v1 Announce Type: new Abstract: Large Language Models (LLMs) show promise as tool-using agents but remain limited in long-horizon tasks that require remembering, organizing, and reusin

Apple unveils an Apple Intelligence feature to automatically change compromised passwords, using agentic AI, and save them to the Passwords app (James Pero/Gizmodo)

AgentsDGX agent

James Pero / Gizmodo: Apple unveils an Apple Intelligence feature to automatically change compromised passwords, using agentic AI, and save them to the Passwords app — Agentic AI and security are norm

At @tryramp, engineers use Devin Desktop to bring their favorite agents into one place. With Devin Desktop they can dispatch, monitor, and j…

AgentsDGX agent

At @tryramp, engineers use Devin Desktop to bring their favorite agents into one place. With Devin Desktop they can dispatch, monitor, and jump between agents from a single surface with shared context

AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning

AgentsDGX agent

arXiv:2512.13278v2 Announce Type: replace Abstract: Agentic reinforcement learning has advanced large language models (LLMs) to reason through long chain-of-thought trajectories while interleaving ext

Declarative Skills for AI Agents in Knowledge-Grounded Tool-Use Workflows

Model ReleasesDGX agent

arXiv:2606.06923v1 Announce Type: new Abstract: We study orchestration mechanisms for tool-using AI agents in realistic customer-service workflows over an unstructured knowledge base. We argue that de

← Previous
1…4849505152…297
Next →