AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,762 results
8 Jun 2026

The Sim-to-Real Gap of Foundation Model Agents: A Unified MDP Perspective

AgentsDGX agent

arXiv:2606.07017v1 Announce Type: new Abstract: Foundation model agents are increasingly deployed for real-world decision-making, but suffer from the sim-to-real gap. While robotics and classical cont

6 Jun 2026

AdaMEM: Test-Time Adaptive Memory for Language Agents

SafetyDGX agent

arXiv:2606.05684v1 Announce Type: new Abstract: A central challenge for language agents is utilizing past experience to adapt to dynamic test-time conditions. While recent work demonstrates the promis

Beyond Semantic Organization: Memory as Execution State Management for Long-Horizon Agents

AgentsDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.06090v1 Announce Type: new Abstract: LLM-based agents increasingly tackle long-horizon tasks with interdependent decisions, where each action reshapes future constraints and intermediate er

“de log is de agent”… activegraph for EU regulation? (just found this) https://djimit.nl/blog/activegraph-event-sourced-agents

AgentsDGX agent

This post explores the concept of using ActiveGraph with event sourcing for AI agents operating under EU regulations, suggesting a potential architecture where detailed logging and event histories ser

Insurance of Agentic AI

AgentsDGX agent

arXiv:2606.05449v1 Announce Type: new Abstract: Agentic artificial intelligence (AI) systems are transforming the risk landscape by extending beyond information generation to autonomous planning, tool

Knowledge Activation: AI Skills as the Institutional Knowledge Primitive for Agentic Software Development

AgentsDGX agent

arXiv:2603.14805v2 Announce Type: replace Abstract: Enterprise software organizations accumulate critical institutional knowledge - architectural decisions, deployment procedures, compliance policies,

5 Jun 2026

Personal AI Agent for Camera Roll VQA

AgentsDGX agent

arXiv:2606.05275v1 Announce Type: new Abstract: We study the personal camera roll visual question answering setting. In this setting, a conversational AI assistant can access a user's personal camera

Retrospective Harness Optimization: Improving LLM Agents via Self-Preference over Trajectory Rollouts

AgentsDGX agent

arXiv:2606.05922v1 Announce Type: cross Abstract: AI agents rely on a harness of skills, tools, and workflows to solve complex problems. Continually improving this harness is essential for adapting to

4 Jun 2026

AIP: A Graph Representation for Learning and Governing Agent Skills

Model ReleasesDGX agent

arXiv:2606.04781v1 Announce Type: new Abstract: Agent Skills today consist largely of free-form prose requiring the agent to read, interpret, and re-derive how to act in every session. This imposes tw

Beyond Prompt-Based Planning: MCP-Native Graph Planning-based Biomedical Agent System

AgentsDGX agent

arXiv:2606.04494v1 Announce Type: new Abstract: Biomedical agents promise to automate complex biological workflows, yet current systems face two fundamental bottlenecks: bioinformatics tools are highl

Exploring the Topology and Memory of Consensus: How LLM Agents Agree, Fragment, or Settle When Forming Conventions

Model ReleasesDGX agent

arXiv:2606.04197v1 Announce Type: cross Abstract: How much should an LLM agent remember, and how should multi-agent systems be connected when trying to reach consensus? We show these two design choice

I am hooked on Dynamic Workflows! The idea of generating harnesses on the fly is so compelling that I reverse-engineered it for my agent orc…

Model ReleasesDGX agent

I am hooked on Dynamic Workflows! The idea of generating harnesses on the fly is so compelling that I reverse-engineered it for my agent orchestrator. And then I built a monitoring dashboard (as an HT

its been such fun befriending Pari and seeing him completely reinvent his company for the agentic era, WHILE having the most insanely stacke…

AgentsDGX agent

its been such fun befriending Pari and seeing him completely reinvent his company for the agentic era, WHILE having the most insanely stacked customer base I've ever seen in the hardest engineering do

MIRAGE: Mobile Agents with Implicit Reasoning and Generative World Models

AgentsDGX agent

arXiv:2606.04627v1 Announce Type: new Abstract: Mobile agents are increasingly expected to operate everyday applications from screenshots and language goals, where reliable control requires reasoning

Strabo: Declarative Specification and Implementation of Agentic Interaction Protocols

AgentsDGX agent

arXiv:2606.05043v1 Announce Type: new Abstract: The last few years have witnessed major advances in the modeling and implementation of multiagent systems based on declarative interaction protocols. Ou

Topology Matters: Measuring Memory Leakage in Multi-Agent LLMs

AgentsDGX agent

arXiv:2512.04668v4 Announce Type: replace-cross Abstract: Graph topology is a fundamental determinant of memory leakage in multi-agent LLM systems, yet its effects remain poorly quantified. We introdu

3 Jun 2026

Automata-Conditioned Cooperative Multi-Agent Reinforcement Learning

AgentsDGX agent

arXiv:2511.02304v2 Announce Type: replace-cross Abstract: We study learning multi-task, multi-agent policies for cooperative, temporal objectives, under centralized training, decentralized execution.

Chopped it up with @swyx on @latentspacepod and we ran the gamut on this one. We talked platform, how roles are evolving, the agentic era, t…

AgentsDGX agent

Chopped it up with @swyx on @latentspacepod and we ran the gamut on this one. We talked platform, how roles are evolving, the agentic era, the future of open source, and what we’re building next. Spoi

RealClawBench: Live OpenClaw Benchmarks from Real Developer-Agent Sessions

Model ReleasesDGX agent

arXiv:2606.03889v1 Announce Type: new Abstract: Agent benchmarks should reflect what users actually ask deployed agents to do, yet existing benchmarks often miss key realism properties of real develop

Sema4.ai’s autonomous agent-building platform gets simpler to use, adds deeper business context and more

AgentsDGX agent

Sema4.ai Inc. a startup that provides tools for building and managing artificial intelligence agents, today announced a massive revamp of its platform, with big changes coming to every layer of the ag

ToolGate: Token-Efficient Pre-Call Control for Tool-Augmented Vision-Language Agents

AgentsDGX agent

arXiv:2606.03054v1 Announce Type: new Abstract: Tool-augmented vision-language agents can acquire external perceptual evidence through OCR, detection, segmentation, and other tools, but executing ever

WRIT: Write-Read Intensive Trajectory Synthesis for Multi-Turn User-Facing Agents

Model ReleasesDGX agent

arXiv:2606.02908v1 Announce Type: cross Abstract: Multi-turn user-facing agents must infer user intent from incomplete requests, collect missing information through dialogue and tools, and execute val

2 Jun 2026

Agentic Authoring of Interactive Multiview Visualizations in Genomics

AgentsDGX agent

arXiv:2606.00370v1 Announce Type: cross Abstract: Diverse genomics data, scientific questions, and analysis tasks typically demand highly specialized visualizations. Therefore, users often must custom

AgentxGCore: Agentic AI for Next-Generation Mobile Core Network

AgentsDGX agent

arXiv:2606.00417v1 Announce Type: cross Abstract: To meet the stringent requirements of emerging applications and the increasingly complex network management and operation, the Next Generation Mobile

Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?

AgentsDGX agent

arXiv:2603.03202v3 Announce Type: replace Abstract: As large language models (LLMs) advance their mathematical capabilities toward the IMO and research level, the scarcity of challenging, high-quality

Codex CLI now runs inside Devin Desktop. You can pair your Codex subscription with Devin Desktop to have native multi-agent orchestration. L…

AgentsDGX agent

Codex CLI now runs inside Devin Desktop. You can pair your Codex subscription with Devin Desktop to have native multi-agent orchestration. Learn more about using your favorite agents in Devin Desktop:

Critic-R: Improving Agentic Search using Instruction-tuned Retrievers with Natural Language Introspective Feedback

AgentsDGX agent

arXiv:2606.00590v1 Announce Type: cross Abstract: Agentic search systems iteratively interact with retrieval models to answer complex queries. Despite substantial progress, optimizing retrievers for a

Defenses & Enablers For Skill Injection Attacks on Terminal Based Agents

AgentsDGX agent

arXiv:2606.01567v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly rely on reusable skills i.e. documents describing task-specific procedures. However, this introduces a

Do Multimodal Agents Really Benefit from Tool Use? A Systematic Study of Capability Gains

Model ReleasesDGX agent

arXiv:2606.02357v1 Announce Type: cross Abstract: Tool-augmented multimodal agents show strong benchmark gains, often taken as evidence that agents have learned to use tools. We argue that this interp

Ev-Trust: An Evolutionarily Stable Trust Mechanism for Decentralized LLM-Based Multi-Agent Service Economies

AgentsDGX agent

arXiv:2512.16167v3 Announce Type: replace-cross Abstract: Decentralized LLM-based multi-agent service economies face three vulnerabilities that undermine traditional trust mechanisms: reduced cost of

Introducing Devin Desktop: the next generation of Windsurf Manage fleets of local and cloud agents from one surface Support for any ACP-comp…

AgentsDGX agent

Introducing Devin Desktop: the next generation of Windsurf Manage fleets of local and cloud agents from one surface Support for any ACP-compatible agent With a full IDE for when you need to jump into

Microsoft’s Project Solara is an OS for AI agent gadgets

AgentsDGX agent

Microsoft just announced 'Project Solara,' a new OS designed for gadgets that run AI agents, at Build 2026. The company is calling it 'a new platform built from the ground up to power agent-driven exp

MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills?

Model ReleasesDGX agent

arXiv:2606.01993v1 Announce Type: cross Abstract: Abundant procedural knowledge on the Web holds great potential for helping agents solve long-horizon tasks. However, such knowledge is often multimoda

ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use

Model ReleasesDGX agent

arXiv:2606.00341v1 Announce Type: cross Abstract: As AI agents are increasingly deployed in real personal and corporate settings (email accounts, development workflows, company databases, etc.), safet

TechGraphRAG: An Agentic Graph-Augmented RAG Framework for Technical Literature Reasoning

AgentsDGX agent

arXiv:2606.01613v1 Announce Type: cross Abstract: This paper presents an agentic retrieval-augmented generation (RAG) framework for domain-specific technical reasoning support, instantiated over a cur

When Parallelism Pays Off: Cohesion-Aware Task Partitioning for Multi-Agent Coding

Model ReleasesDGX agent

arXiv:2606.00953v1 Announce Type: new Abstract: Multi-agent Large Language Model (LLM) systems offer a way to decompose complex tasks, such as coding, through parallelization and context isolation. Ho

1 Jun 2026

Beyond LLMs: Why Scalable Enterprise AI Adoption Depends on Agent Logic

AgentsDGX agent

Enterprise AI adoption at scale requires moving beyond large language models to implement agent logic systems that can handle complex reasoning, planning, and decision-making autonomously. Agent-based

Connecting agents directly to Snowflake, Slack, or BigQuery via MCP/CLIs is a multi-token disaster. They default to brute-force exploration,…

AgentsDGX agent

Connecting agents directly to Snowflake, Slack, or BigQuery via MCP/CLIs is a multi-token disaster. They default to brute-force exploration, firing 20-30 tool calls just to rediscover context on every

Counterfactual Trace Auditing of LLM Agent Skills

Model ReleasesDGX agent

arXiv:2605.11946v2 Announce Type: replace Abstract: Large Language Model agents are increasingly augmented with agent skills. Current evaluation methods for skills remain limited. Most deployed benchm

DynaTree: Dynamic Agentic Retrieval Tree for Time-Sensitive News Retrieval

Model ReleasesDGX agent

arXiv:2605.31377v1 Announce Type: cross Abstract: Agentic Retrieval-Augmented Generation improves retrieval by integrating planning, tool use, and iterative reasoning, but existing agentic RAG methods

if you build any agent on activegraph, the trace is automatic and first-class, not bolted on

AgentsDGX agent

if you build any agent on activegraph, the trace is automatic and first-class, not bolted on a parallel experiment building a coding agent on top of @activegraphai. you can see everything flattened do

LLM Anonymization Against Agentic Re-Identificatio

AgentsDGX agent

arXiv:2605.30848v1 Announce Type: cross Abstract: Agentic LLMs with web search change the threat model for text anonymization: weak contextual cues can become cross-referenceable evidence for re-ident

More info about Search as Code in the Perplexity Agent API docs: https://docs.perplexity.ai/docs/agent-api/tools/sandbox

AgentsDGX agent

The Perplexity Agent API documentation includes a 'Search as Code' feature accessible through the sandbox tools section, enabling developers to integrate search functionality programmatically within a

31 May 2026

🧑‍⚖️Evaluating Deep Agents with LangSmith on AWS Great deep dive blog with our friends at AWS on evaluating DeepAgents with LangSmith Cover…

AgentsDGX agent

🧑‍⚖️Evaluating Deep Agents with LangSmith on AWS Great deep dive blog with our friends at AWS on evaluating DeepAgents with LangSmith Covers datapoint and evaluator design for longer horizon agents ht

30 May 2026

Learn more about the latest from @james_y_zou and our Frontier Agents Research team!

AgentsDGX agent

Learn more about the latest from @james_y_zou and our Frontier Agents Research team! To evaluate frontier AI agents, we need more complex tasks. But such tasks are also more prone to have design mista

Step 3.7 Flash is now free for 30 days via Nous Portal It is a new MoE vision-language model focused on agent efficiency, coding, search, an…

AgentsDGX agent

Step 3.7 Flash is now free for 30 days via Nous Portal It is a new MoE vision-language model focused on agent efficiency, coding, search, and multimodal workflows — and Hermes Agent users have been lo

29 May 2026

AgentSchool: An LLM-Powered Multi-Agent Simulation for Education

AgentsDGX agent

arXiv:2605.30144v1 Announce Type: new Abstract: Despite the rapid deployment of LLMs into classrooms, validating educational AI remains uniquely intractable: interventions act on developing learners w

Does The Way You Plan Matter? An Empirical Study of Planning Representations for LLM Web Agents

AgentsDGX agent

arXiv:2605.29927v1 Announce Type: cross Abstract: Despite recent advances, LLM-based web agents still struggle with limited exploration, omission of critical steps, and sensitivity to task constraints

Governing Technical Debt in Agentic AI Systems

AgentsDGX agent

arXiv:2605.29129v1 Announce Type: new Abstract: Agentic AI systems are increasingly being explored as production infrastructure: they reason over multiple steps, call tools, act through workflows, and

Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction

AgentsDGX agent

arXiv:2605.29960v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly leverage long term memory to support persistent and autonomous task execution. However, this capability

make something agents want

AgentsDGX agent

make something agents want studying ActiveGraph by @yoheinakajima and apart from the concept the implementation itself is brilliant on the site it shares a prompt that will make my agent study docs, i

OpenClawBench: Benchmarking Process-side Anomalies in Real-world Agent Execution Trajectories

Model ReleasesDGX agent

arXiv:2605.29253v1 Announce Type: new Abstract: Task success can hide process anomalies in real-world agent executions. An agent may pass the final task oracle while still accumulating unresolved ambi

Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generation

AgentsDGX agent

arXiv:2605.29861v1 Announce Type: cross Abstract: Large Language Models (LLMs) have advanced autonomous agents from deep search, which retrieves concise factual answers, to deep research, which synthe

28 May 2026

a hot (cold at this point?) take that lead us to build this: every agent in the future will need a sandbox to connect to writing/executing c…

AgentsDGX agent

a hot (cold at this point?) take that lead us to build this: every agent in the future will need a sandbox to connect to writing/executing code is not just for coding agents! is useful for all sorts o

A Unified Framework for the Evaluation of LLM Agentic Capabilities

Model ReleasesDGX agent

arXiv:2605.27898v1 Announce Type: new Abstract: As LLMs are increasingly deployed as agents, reliable assessment of their agentic capabilities has become essential. However, reported benchmark scores

Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents

AgentsDGX agent

arXiv:2605.22166v2 Announce Type: replace Abstract: LLM agents are shaped not only by their language models, but also by the runtime harness that mediates observation, tool use, action execution, feed

Ask Now, Use Later: Benchmarking the Proactivity Gap in Long-Lived LLM Agents

Model ReleasesDGX agent

arXiv:2605.28108v1 Announce Type: new Abstract: A long-lived LLM agent, such as OpenClaw, earns its value by acting on a user's preferences and constraints across sessions, not just the current reques

EgoBench: An Interactive Egocentric Multimodal Benchmark for Tool-Using Agents

Model ReleasesDGX agent

arXiv:2605.27820v1 Announce Type: new Abstract: As AI agents increasingly operate in open, real-world environments, they require a deep synergy of multimodal perception, tool invocation with multi-hop

From Instructor to Collaborator: What a 90-Participant Study Reveals about Human-Agent Collaboration in a Mobile Serious Game

AgentsDGX agent

arXiv:2605.27384v1 Announce Type: cross Abstract: This position paper reflects empirical data collected during my PhD from a large-scale within-subjects study (N = 90). The study compared a highly hum

GUI-CIDER: Mid-training GUI Agents via Causal Internalization and Density-aware Exemplar Reselection

AgentsDGX agent

arXiv:2605.28534v1 Announce Type: new Abstract: Despite the rapid progress of multimodal large language models in building Graphical User Interface (GUI) agents, their real-world task completion is fu

← Previous
1…4950515253…297
Next →