AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,771 results
22 May 2026

Superhuman Safe and Agile Racing through Multi-Agent Reinforcement Learning

SafetyDGX agent

arXiv:2605.22748v1 Announce Type: new Abstract: Autonomous systems have achieved superhuman performance in isolation or simulation, yet they remain brittle in shared, dynamic real-world spaces. This f

these guys built a research agent with activegraph (using their @monid_ai tool) and found that every claim was traced to a source - which is…

AgentsDGX agent

these guys built a research agent with activegraph (using their @monid_ai tool) and found that every claim was traced to a source - which is not prompted, but natively baked in to the approach (they a

21 May 2026

Giving Agents Computers — Ivan Burazin, Daytona

Agents
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

Latent Space episode featuring Ivan Burazin discussing Daytona, a platform or framework that enables AI agents to use computers and interact with software systems autonomously. The discussion likely c

Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents

AgentsDGX agent

arXiv:2605.21347v1 Announce Type: cross Abstract: Diagnosing failures in LLM agents remains largely manual. Practitioners inspect a small subset of execution traces, form ad-hoc hypotheses, and iterat

Introducing the sandbox Auth Proxy: A way to control the boundary between agent-generated behavior and the rest of the world. An explainer f…

AgentsDGX agent

The sandbox Auth Proxy is a security mechanism designed to control and limit the interactions between AI agents and external systems by establishing a controlled boundary. This explainer from Harrison

Reimagining ML Operations with Agent Skills: a new maturity model for on-call

AgentsDGX agent

This article presents a maturity model for ML operations that leverages agent skills to improve on-call practices and incident response workflows. It likely discusses how organizations can evolve thei

streaming from modern agents is pretty complex! especially with - parallel tools / subagents - multimodal content - human in the loop events…

AgentsDGX agent

streaming from modern agents is pretty complex! especially with - parallel tools / subagents - multimodal content - human in the loop events our new streaming primitives make all of this ergonomic! Ag

20 May 2026

AQuaUI: Visual Token Reduction for GUI Agents with Adaptive Quadtrees

AgentsDGX agent

arXiv:2605.19260v1 Announce Type: new Abstract: Large Multimodal Models (LMMs) have recently emerged as promising backbones for GUI-agent models, where high-resolution GUI screenshots are introduced t

datasette-agent-charts 0.1a1

AgentsDGX agent

Release: datasette-agent-charts 0.1a1 More color! Bar and waffle charts without a color column are shaded by magnitude with a sequential color scheme; color columns holding text values use the observa

From performance reviews to pink slips, managing AI agents looks a lot like managing people

AgentsDGX agent

The workforce is no longer purely human — and closing the gap between how companies manage people and how they govern AI agents has become one of the defining operational challenges of the digital wor

i'm excited to open source Active Graph: an event-sourced reactive graph runtime for long-running, agents 🔄🧠 events/logs projects a graph.…

AgentsDGX agent

i'm excited to open source Active Graph: an event-sourced reactive graph runtime for long-running, agents 🔄🧠 events/logs projects a graph. reactive behaviors react and affect the graph. fork-and-diff

Love this concept. Always been a proponent of event-driven architectures, and the actor model. Neat to see how it’s applied to the agent run…

AgentsDGX agent

Love this concept. Always been a proponent of event-driven architectures, and the actor model. Neat to see how it’s applied to the agent runtime here. i'm excited to open source Active Graph: an event

Metric-Gradient Projection for Stable Multi-Agent Policy Learning

SafetyDGX agent

arXiv:2605.18809v1 Announce Type: cross Abstract: General-sum multi-agent learning is often governed by a stacked update field in which each agent's policy update changes the optimization landscape fa

Search Self-play: Pushing the Frontier of Agent Capability without Supervision

Model ReleasesDGX agent

arXiv:2510.18821v3 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) has become the mainstream technique for training LLM agents. However, RLVR highly depends on w

Terra Security unifies web, AI and network testing under one agentic platform

AgentsDGX agent

Agentic offensive security platform provider Terra Security Inc. today announced the launch of continuous exploitation validation for network infrastructure, extending the company’s platform beyond we

there he goes again, @yoheinakajima blazing the paths to continuously running self aware agents. this and the posts he links are super worth…

AgentsDGX agent

there he goes again, @yoheinakajima blazing the paths to continuously running self aware agents. this and the posts he links are super worth reading. i'm excited to open source Active Graph: an event-

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?

Model ReleasesDGX agent

arXiv:2605.19196v1 Announce Type: new Abstract: Deep research agents increasingly automate complex information-seeking tasks, producing evidence-grounded reports via multi-step reasoning, tool use, an

Towards Trust Calibration in Socially Interactive Agents: Investigating Gendered Multimodal Behaviors Generation with LLMs

Model ReleasesDGX agent

arXiv:2605.19798v1 Announce Type: new Abstract: As Socially Interactive Agents (SIAs) become increasingly integrated into daily life, the ability to calibrate user trust to an agent's actual capabilit

Vision Harnessing Agent for Open Ad-hoc Segmentation

Model ReleasesDGX agent

arXiv:2605.19410v1 Announce Type: new Abstract: Segmentation has become easy when the concept is known, requiring retrieval of a learned visual grounding from text. It remains hard for open ad-hoc con

We built an AI agent for due diligence, with exact audit trails back to the source page, that you can use as a template without paying a sin…

AgentsDGX agent

We built an AI agent for due diligence, with exact audit trails back to the source page, that you can use as a template without paying a single dime for PDF parsing 🔥🆓 The secret sauce is LiteParse -

We ran 720 browser agent tasks with @nottecore across frontier models. One baseline model produced malformed outputs in ~1 out of every 5 ca…

AgentsDGX agent

We ran 720 browser agent tasks with @nottecore across frontier models. One baseline model produced malformed outputs in ~1 out of every 5 calls, leading to retries inside multi-step workflows. Across

19 May 2026

1GC-7RC: One Graphic Card -- Seven Research Challenges! How Good Are AI Agents at Doing Your Job?

Model ReleasesDGX agent

arXiv:2605.17046v1 Announce Type: cross Abstract: Autonomous AI coding agents are becoming a core tool for ML practitioners in industry and research alike. Despite this growing adoption, no standardiz

ASPI: Seeking Ambiguity Clarification Amplifies Prompt Injection Vulnerability in LLM Agents

Model ReleasesDGX agent

arXiv:2605.17324v1 Announce Type: cross Abstract: Clarification-seeking behavior is widely regarded as a desirable property of LLM agents, enabling them to resolve ambiguity before acting on underspec

Databricks context engineer associate: the industry’s first certification for reliable AI agent systems

AgentsDGX agent

Databricks has launched the Context Engineer Associate certification, described as the industry's first credential specifically designed for professionals building reliable AI agent systems. The certi

From Prompts to Protocols: An AI Agent for Laboratory Automation

AgentsDGX agent

arXiv:2605.16552v1 Announce Type: new Abstract: Automating science laboratories enables faster, safer, more accurate, and more reproducible execution of protocols, accelerating the discovery and testi

Heterogeneous Information-Bottleneck Coordination Graphs for Multi-Agent Reinforcement Learning

AgentsDGX agent

arXiv:2605.17393v1 Announce Type: new Abstract: Coordination graphs are a central abstraction in cooperative multi-agent reinforcement learning (MARL), yet existing sparse-graph learners lack a theore

Lying with Truths: Open-Channel Multi-Agent Collusion for Belief Manipulation via Generative Montage

AgentsDGX agent

arXiv:2601.01685v2 Announce Type: replace-cross Abstract: As large language models (LLMs) transition to autonomous agents synthesizing real-time information, their reasoning capabilities introduce an

Overeager Coding Agents: Measuring Out-of-Scope Actions on Benign Tasks

Model ReleasesDGX agent

arXiv:2605.18583v1 Announce Type: cross Abstract: Coding agents now run autonomously with shell, file, and network privileges. When a user issues a benign request, the agent sometimes does more than a

SEMA-RAG: A Self-Evolving Multi-Agent Retrieval-Augmented Generation Framework for Medical Reasoning

AgentsDGX agent

arXiv:2605.17101v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) is widely employed to mitigate risks such as hallucinations and knowledge obsolescence in medical question answer

The Alpha Illusion: Reported Alpha from LLM Trading Agents Should Not Be Treated as Deployment Evidence

AgentsDGX agent

arXiv:2605.16895v1 Announce Type: cross Abstract: End-to-end LLM trading agents have moved quickly from research curiosity to a small ecosystem of named systems, including FinCon, FinMem, TradingAgent

Voker raises $2.2M to help teams understand how AI agents perform in the wild

AgentsDGX agent

Voker, an agent analytics platform for artificial intelligence product teams, today announced it has raised 2.2 million in pre-seed funding from Y Combinator and FundersClub. As more companies push AI

18 May 2026

AgentStop: Terminating Local AI Agents Early to Save Energy in Consumer Devices

Local AiDGX agent

arXiv:2605.15206v1 Announce Type: cross Abstract: Autonomous agents powered by large language models (LLMs) are increasingly used to automate complex, multi-step tasks such as coding or web-based ques

Ansible positions automation as the trusted execution layer for agentic AI

AgentsDGX agent

Ansible, Red Hat Inc.’s automation platform, is emerging as the trusted execution layer that might bridge AI-generated insights and reliable IT operations, aiming to turn probabilistic agentic AI into

Is Agentic AI Ready for Real-World Hardware Engineering? A Deep Dive with Phoenix-bench

Local AiDGX agent

arXiv:2605.15226v1 Announce Type: cross Abstract: We ask whether agentic AI systems built for software engineering transfer to realistic hardware engineering. Existing hardware LLM benchmarks isolate

SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization

SafetyDGX agent

arXiv:2604.02268v2 Announce Type: replace Abstract: Agent skills, structured packages of procedural knowledge and executable resources that agents dynamically load at inference time, have become a rel

Traj-CoA: Patient Trajectory Modeling via Chain-of-Agents for Lung Cancer Risk Prediction

AgentsDGX agent

arXiv:2510.10454v2 Announce Type: replace Abstract: Large language models (LLMs) offer a generalizable approach for modeling patient trajectories, but suffer from the long and noisy nature of electron

17 May 2026

WebWright - Agentic Extension for your Browser is now live on Chrome 😁

AgentsDGX agent

WebWright is a Chrome browser extension that enables AI agents to automate web tasks directly within the browser, leveraging local models like those available through Ollama. The announcement on r/oll

15 May 2026

A Two-Dimensional Framework for AI Agent Design Patterns: Cognitive Function and Execution Topology

AgentsDGX agent

arXiv:2605.13850v1 Announce Type: new Abstract: Existing frameworks for LLM-based agent architectures describe systems from a single perspective: industry guides (Anthropic, Google, LangChain) focus o

Agents are a manager’s dream for productivity — and a CISO’s worst nightmare when they go rogue

AgentsDGX agent

Agent risk management is evolving rapidly as AI moves into consequential decision-making arenas, collapsing the boundary between human and machine risk across the enterprise. The dual-threat landscape

ChromaFlow: A Negative Ablation Study of Orchestration Overhead in Tool-Augmented Agent Evaluation

AgentsDGX agent

arXiv:2605.14102v1 Announce Type: new Abstract: Autonomous language-model agents increasingly combine planning, tool use, document processing, browsing, code execution, and verification loops. These c

Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction

Model ReleasesDGX agent

arXiv:2605.13950v1 Announce Type: cross Abstract: Autonomous language-model agents are increasingly evaluated on long-horizon tool-use tasks, but existing benchmarks rarely capture the complexity and

Excited to see SmithDB announcement at Interrupt, our purpose-built distributed database for agent observability! SmithDB is built on top ob…

AgentsDGX agent

Excited to see SmithDB announcement at Interrupt, our purpose-built distributed database for agent observability! SmithDB is built on top object storage, written in Rust and leverages Apache DataFusio

Great paper discussing agentic search vs. vector search.

AgentsDGX agent

Great paper discussing agentic search vs. vector search. // Is Grep All You Need? // Pay attention to this on, AI devs. (bookmark it) They find that grep-style text search, when wrapped in the right a

Invisible Orchestrators Suppress Protective Behavior and Dissociate Power-Holders: Safety Risks in Multi-Agent LLM Systems

Model ReleasesDGX agent

arXiv:2605.13851v1 Announce Type: new Abstract: Multi-agent orchestration -- in which a hidden coordinator manages specialized worker agents -- is becoming the default architecture for enterprise AI d

MALLVI: A Multi-Agent Framework for Integrated Generalized Robotics Manipulation

AgentsDGX agent

arXiv:2602.16898v5 Announce Type: replace-cross Abstract: Task planning for robotic manipulation with large language models (LLMs) is an emerging area. Prior approaches rely on specialized models, fin

Near-Miss: Latent Policy Failure Detection in Agentic Workflows

Model ReleasesDGX agent

arXiv:2603.29665v2 Announce Type: replace Abstract: Agentic systems for business process automation often require compliance with policies governing conditional updates to the system state. Evaluation

PREPING: Building Agent Memory without Tasks

AgentsDGX agent

arXiv:2605.13880v1 Announce Type: new Abstract: Agent memory is typically constructed either offline from curated demonstrations or online from post-deployment interactions. However, regardless of how

Ready from Day 1: Population-Aware Coordination for Large-Scale Constrained Multi-Agent Systems

AgentsDGX agent

arXiv:2605.13900v1 Announce Type: cross Abstract: In large-scale multi-agent systems with shared resource constraints, an upstream planner must iteratively evaluate candidate resource plans -- assessi

SimPersona: Learning Discrete Buyer Personas from Raw Clickstreams for Grounded E-Commerce Agents

SafetyDGX agent

arXiv:2605.14205v1 Announce Type: new Abstract: LLM-based web agents can navigate live storefronts, yet they often collapse to a single 'average buyer' policy, failing to capture the heterogeneous and

SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades

Model ReleasesDGX agent

arXiv:2605.14415v1 Announce Type: cross Abstract: Coding agents powered by large language models are increasingly expected to perform realistic software maintenance tasks beyond isolated issue resolut

Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining

AgentsDGX agent

arXiv:2605.14747v1 Announce Type: cross Abstract: Recent advances in multimodal large language models have driven growing interest in graphical user interface (GUI) agents, yet their generalization re

Why Neighborhoods Matter: Traversal Context and Provenance in Agentic GraphRAG

AgentsDGX agent

arXiv:2605.15109v1 Announce Type: new Abstract: Retrieval-Augmented Generation can improve factuality by grounding answers in external evidence, but Agentic GraphRAG complicates what it means for cita

14 May 2026

Building Interactive Real-Time Agents with Asynchronous I/O and Speculative Tool Calling

Model ReleasesDGX agent

arXiv:2605.13360v1 Announce Type: new Abstract: There is a growing demand for agentic AI technologies for a range of downstream applications like customer service and personal assistants. For applicat

deepagents v0.6 is our biggest release yet!!! it’s all about perf - at the model layer w harness profiles, agent layer w code interpreter, a…

AgentsDGX agent

deepagents v0.6 is our biggest release yet!!! it’s all about perf - at the model layer w harness profiles, agent layer w code interpreter, and at scale w streaming and delta channels context hub backe

IdeaForge: A Knowledge Graph-Grounded Multi-Agent Framework for Cross-Methodology Innovation Analysis and Patent Claim Generation

AgentsDGX agent

arXiv:2605.13311v1 Announce Type: new Abstract: Current AI-assisted innovation systems typically apply a single ideation methodology (such as TRIZ or Design Thinking) using sequential prompt-based wor

Interesting position paper on agentic AI as a foreseeable pathway to AGI. (bookmark it) There has been strong debate on whether a larger sin…

SafetyDGX agent

Interesting position paper on agentic AI as a foreseeable pathway to AGI. (bookmark it) There has been strong debate on whether a larger single model get us there or a multi-agent system. The authors

Position: Agentic AI System Is a Foreseeable Pathway to AGI

AgentsDGX agent

arXiv:2605.12966v1 Announce Type: new Abstract: Is monolithic scaling the only path to AGI? This paper challenges the dogma that purely scaling a single model is sufficient to achieve Artificial Gener

Sea's View on the Future of Agentic Software Development with Codex

AgentsDGX agent

This article presents Sea's perspective on how agentic software development—systems that can autonomously plan and execute coding tasks—will evolve with tools like Codex, OpenAI's code generation mode

You can now power your Hermes Agent, if using OpenAI models, with codex as the runtime for the core tools that it offers, with the flip of a…

AgentsDGX agent

Nous Research announced that Hermes Agents powered by OpenAI models can now use Codex as the runtime for executing core tools, enabling improved tool execution capabilities with a simple configuration

13 May 2026

A new Database product based on @ApacheDataFusio was announced today from @LangChain -- focused on agent observability. It is really neat to…

AgentsDGX agent

A new Database product based on @ApacheDataFusio was announced today from @LangChain -- focused on agent observability. It is really neat to see how people are building (very) customized data + query

← Previous
1…5960616263…297
Next →