AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,911 results
4 Jun 2026

Offroad launches with $7M to automate identity security with AI agents

Model ReleasesDGX agent

Offroad Inc. launched today with 7 million in funding to build what it calls an agentic identity security team, using artificial intelligence agents to investigate and remediate access risks across hu

Plan First, Judge Later, Run Better: A DMAIC-Inspired Agentic System for Industrial Anomaly Detection

SafetyDGX agent

arXiv:2606.04599v1 Announce Type: new Abstract: Large language model (LLM) agents have shown promise in automating complex data-analysis workflows, but their reliable deployment remains challenging in

Proud to announce that Cohere has been awarded first place in NATO’s Agentic AI for Cognitive Warfare Innovation Challenge. Congratulations …

AgentsDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Proud to announce that Cohere has been awarded first place in NATO’s Agentic AI for Cognitive Warfare Innovation Challenge. Congratulations to our fellow finalists: OpenMinds, which secured second pla

Scaling Datasets for Multi-Sensor, Multi-Agent, and Multi-Domain Learning in Autonomous Systems

AgentsDGX agent

arXiv:2606.04444v1 Announce Type: cross Abstract: Existing datasets cannot support large-scale learning in multi-agent, multi-sensor, or multi-domain autonomy, where diversity and coordination are ess

Self-Reflective APIs: Structure Beats Verbosity for AI Agent Recovery

AgentsDGX agent

arXiv:2606.05037v1 Announce Type: cross Abstract: When an AI agent calls an API and hits a validation error, it needs more than what went wrong -- it needs what to do next. A self-reflective API retur

What If Prompt Injection Never Left? Exploring Cross-Session Stored Prompt Injection in Agentic Systems

Model ReleasesDGX agent

arXiv:2606.04425v1 Announce Type: cross Abstract: Modern agentic systems transform LLMs from session-bounded assistants into stateful systems that persist and evolve shared world state across sessions

3 Jun 2026

A New Framework for Cybersecurity Refusals in AI Agents

Model ReleasesDGX agent

arXiv:2606.02644v1 Announce Type: cross Abstract: Agentic scaffolds have dramatically improved LLM performance on complex, long-horizon tasks, yielding both broad benefits and amplified risks in domai

At @harvey, the engineering team integrated Spectre — their internal background agent — into Devin Desktop. Now Spectre's organizational con…

AgentsDGX agent

At @harvey, the engineering team integrated Spectre — their internal background agent — into Devin Desktop. Now Spectre's organizational context can live on every engineer's laptop and flow across the

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning

AgentsDGX agent

arXiv:2606.03108v1 Announce Type: new Abstract: Autonomous LLM training is often framed as recipe search, which leaves the training harness largely static. This limitation sharpens in agentic RL, wher

Having trouble connecting Hermes Agent Desktop to your remote instance? Check out this updated guide: https://hermes-agent.nousresearch.com/…

AgentsDGX agent

This guide addresses connectivity issues between Hermes Agent Desktop and remote instances, providing updated troubleshooting instructions from Nous Research. It likely covers common connection proble

II-Browser is here: a Chrome extension that lets II-Agent work inside your browser, across the tabs, sessions, and tools you already use. Th…

AgentsDGX agent

II-Browser is here: a Chrome extension that lets II-Agent work inside your browser, across the tabs, sessions, and tools you already use. The browser is no longer just where you work. It’s where your

Inducing Reasoning Primitives from Agent Traces

AgentsDGX agent

arXiv:2606.02994v1 Announce Type: new Abstract: ReAct-style LLM agents often rediscover the same reasoning routines across problems, yet leave those routines trapped in transient scratchpads. We intro

Internal documents: Meta is considering tiered pricing for 'Hatch', its planned OpenClaw-like AI agent tool, including a $200-per-month premium subscription (Jyoti Mann/The Information)

AgentsDGX agent

Jyoti Mann / The Information: Internal documents: Meta is considering tiered pricing for “Hatch”, its planned OpenClaw-like AI agent tool, including a 200-per-month premium subscription — Meta Platfor

Nous Portal is the simplest way to power your Hermes Agent. Run 'hermes portal' to switch today.

AgentsDGX agent

Nous Portal is a tool designed to streamline the deployment and operation of Hermes Agents, offering a simplified interface for users. The portal can be activated by running the 'hermes portal' comman

On dynamic multi-agent pathfinding methods: review, simulations and modifications

AgentsDGX agent

arXiv:2606.03735v1 Announce Type: cross Abstract: This paper presents a systematic study of pathfinding algorithms in the context of Dynamic Multi-Agent Pathfinding (D-MAPF), a setting that combines d

Open data architecture powers DoorDash’s real-time logistics and agentic AI ambitions

AgentsDGX agent

As enterprises race to support machine learning, agentic workflows and analytics on the same infrastructure, open data architecture has emerged as the dividing line between platforms that scale and th

Perceive Before Reasoning: A Pre-Reasoning Perception Framework for Efficient and Reliable Proactive Mobile Agents

Model ReleasesDGX agent

arXiv:2606.03236v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have substantially advanced mobile agents, yet proactive mobile assistance remains challenging because agents m

The Agent's First Day: Benchmarking Learning, Exploration, and Scheduling in the Workplace Scenarios

AgentsDGX agent

arXiv:2601.08173v2 Announce Type: replace Abstract: The rapid evolution of Multi-modal Large Language Models (MLLMs) has advanced workflow automation; however, existing research mainly targets perform

The Epi-LLM Framework: probing LLM behavioral priors through epidemiological agent-based models

AgentsDGX agent

arXiv:2606.02867v1 Announce Type: cross Abstract: Human behaviour during epidemics affects infectious disease dynamics, but quantifying this remains deeply challenging. Here we introduce the Epi-LLM f

Towards a Science of AI Agent Reliability

SafetyDGX agent

arXiv:2602.16666v3 Announce Type: replace Abstract: AI agents are increasingly deployed to execute important tasks. While rising accuracy scores on standard benchmarks suggest rapid progress, many age

Very cool work from the Quarq team Their agent is the current LongMemEval-S leader, and it's built on LangGraph Many such cases :)

AgentsDGX agent

Quarq's team has developed an AI agent built on LangGraph that currently leads the LongMemEval-S benchmark, demonstrating advances in long-context memory evaluation. The work represents a notable achi

When Helping Hurts and How to Fix It: Multi-Agent Debate for Data Cleaning

AgentsDGX agent

arXiv:2606.02866v1 Announce Type: new Abstract: When does multi-agent debate help data cleaning, and when does it hurt? Across three benchmarks, four model families, and over 6,000 task-condition pair

Whoop is building agentic AI maturity on a foundation of enterprise health data

AgentsDGX agent

Enterprise AI programs are moving beyond experimentation, but agentic AI maturity — the ability to run governed, autonomous workflows at production scale — remains out of reach for most organizations.

2 Jun 2026

Adaptive Auto-Harness: Sustained Self-Improvement for Agentic System Deployment on Open-Ended Task Streams

AgentsDGX agent

arXiv:2606.01770v1 Announce Type: cross Abstract: Auto-harness systems such as A-Evolve, GEPA, and Meta-Harness improve LLM agents by optimizing prompts, skills, tools, memories, and supporting infras

Agentic Transformers Provably Learn to Search via Reinforcement Learning

SafetyDGX agent

arXiv:2606.00183v1 Announce Type: cross Abstract: Tree search is a central abstraction behind many language-agent reasoning and decision-making tasks: agents must explore actions, remember failures, a

AI agents, open data and governance take center stage at Snowflake Summit

AgentsDGX agent

Snowflake Inc. is using its Summit 2026 conference today in San Francisco to present a vision of what it calls the “agentic enterprise,” unveiling a broad set of products and enhancements that it says

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models

Model ReleasesDGX agent

arXiv:2606.01961v1 Announce Type: new Abstract: Autonomous agents are increasingly expected to support end-to-end medical-AI research workflows, moving beyond isolated prediction tasks or short-form c

AWS adds database features and license options aimed at simplifying agent deployment

AgentsDGX agent

Amazon Web Services Inc. today enhanced its database services to simplify the process of building and operating agentic artificial intelligence applications while also lowering the barriers to cloud m

BADGER: Bridging Agentic and Deterministic Evaluation for Generative Enterprise Reasoning

Model ReleasesDGX agent

arXiv:2606.02109v1 Announce Type: new Abstract: Enterprise AI systems that translate natural language into SQL queries and orchestrate multi-step agentic reasoning pipelines require evaluation approac

Beyond End-to-End Video Models: An LLM-Based Multi-Agent System for Educational Video Generation

SafetyDGX agent

arXiv:2602.11790v2 Announce Type: replace Abstract: Although recent end-to-end video generation models demonstrate impressive performance in visually oriented content creation, they remain limited in

BraveGuard: From Open-World Threats to Safer Computer-Use Agents

Model ReleasesDGX agent

arXiv:2606.01166v1 Announce Type: cross Abstract: Computer-use agents extend language models from text generation to sustained interaction with files, terminals, browsers, and external tools. This shi

Context Matters: Repository-Aware Security Analysis of the Agent Skill Ecosystem

Model ReleasesDGX agent

arXiv:2603.16572v2 Announce Type: replace-cross Abstract: Agent skills extend local AI agents, such as Claude Code and OpenClaw, with additional functionality. Their growing popularity has led to dedi

Diversity Over Frequency: Rethinking Tool Use in Visual Chain-of-Thought Agents

AgentsDGX agent

arXiv:2606.00096v1 Announce Type: cross Abstract: Visual agents employ external visual tools within visual chains of thought to incorporate fine-grained evidence. While prior work has mainly studied t

Early research with @LangChain Labs on building more efficient verifiers for agent work product.

AgentsDGX agent

LangChain Labs is conducting research into developing more efficient verification systems for validating the outputs and work products of AI agents. This research likely focuses on improving the speed

Efficient verification is what makes scaling legal agents practical. Excited to partner with @hwchase17 and the @LangChain Labs team on desi…

AgentsDGX agent

Efficient verification is what makes scaling legal agents practical. Excited to partner with @hwchase17 and the @LangChain Labs team on designing efficient verifiers - sharing early results showing op

GitHub unveils a GitHub Copilot desktop app in technical preview, which introduces a new feature called canvases for bidirectional work between users and agents (Mario Rodriguez/The GitHub Blog)

AgentsDGX agent

Mario Rodriguez / The GitHub Blog: GitHub unveils a GitHub Copilot desktop app in technical preview, which introduces a new feature called canvases for bidirectional work between users and agents — At

HLL: Can Agents Cross Humanity's Last Line of Verification?

Model ReleasesDGX agent

arXiv:2606.02449v1 Announce Type: new Abstract: Multimodal agents are increasingly expected to operate interfaces on behalf of users, raising a central deployment question: can they truly substitute f

How Generation Architecture Shapes Code Complexity in Multi-Agent LLM Systems: A Paired Study on HumanEval

AgentsDGX agent

arXiv:2606.00308v1 Announce Type: cross Abstract: Large-language-model code generation has shifted from single-shot prompting to multi-agent orchestrations - analyst, coder, tester, and debugger pipel

Introducing Devin Desktop. Manage fleets of local and cloud agents from one surface. Plan, delegate, review, and ship without leaving your e…

AgentsDGX agent

Devin Desktop is a unified management interface from Cognition AI that enables users to coordinate multiple AI agents operating across local and cloud environments. The platform streamlines workflows

Microsoft debuts an expansion of its model families and agentic AI intelligence for developers

AgentsDGX agent

Microsoft Corp. announced an expansion to its artificial intelligence models and agentic AI infrastructure today that brings more data and context into the hands of developers and business users as th

Network Distributed Multi-Agent Reinforcement Learning for Consensus Control of Quadcopters

SafetyDGX agent

arXiv:2606.02107v1 Announce Type: cross Abstract: This paper proposes a Network Distributed Multi-Agent Reinforcement Learning (ND-MARL) framework for quadcopter consensus control. Compared to convent

Rashomon Memory: Towards Argumentation-Driven Retrieval for Multi-Perspective Agent Memory

AgentsDGX agent

arXiv:2604.03588v3 Announce Type: replace Abstract: AI agents operating over extended time horizons accumulate experiences that serve multiple concurrent goals, and must often maintain conflicting int

Recognize Your Orchestrator: An Entropy Dynamics Perspective for LLM Multi-Agent Systems

AgentsDGX agent

arXiv:2606.01351v1 Announce Type: new Abstract: The transition from single-turn models to Multi-Agent Systems (MAS) promises enhanced problem-solving capabilities, yet the centralized orchestration to

SeClaw: Spec-Driven Security Task Synthesis for Evaluating Autonomous Agents

Model ReleasesDGX agent

arXiv:2606.02302v1 Announce Type: cross Abstract: Autonomous LLM agents increasingly operate in stateful environments where they access tools, files, memory, and external services. While such capabili

Simulating Macroeconomic Expectations in Survey Experiments with LLM-based Economic Agents

ResearchDGX agent

arXiv:2505.17648v5 Announce Type: replace-cross Abstract: We introduce a framework for simulating macroeconomic expectations in survey experiments using LLM-based economic agents (LLM Agents). We cons

SPADE-Bench: Evaluating Spontaneous Strategic Deception in Agents via Plan-Action Divergence

Model ReleasesDGX agent

arXiv:2606.02380v1 Announce Type: cross Abstract: As LLM-based agents expand their operational scope, reliability becomes a prerequisite for real-world deployment. However, in practical applications,

Streaming tokens is now buttery smooth on @telegram with Hermes Agent.

AgentsDGX agent

Nous Research announced improvements to token streaming functionality for the Hermes Agent on Telegram, enhancing the smoothness and performance of real-time token generation. This update likely addre

The agent development lifecycle has been manual for too long. We’re building a future where it runs continuously, without manual triggers. W…

AgentsDGX agent

The agent development lifecycle has been manual for too long. We’re building a future where it runs continuously, without manual triggers. Where well-understood issue types resolve without human revie

The quarq agent is built on LangGraph! LangGraph makes it easy to build complex memory systems (quarq is now at the top of the LongMemEval l…

AgentsDGX agent

The Quarq agent is constructed using LangGraph, a framework that simplifies the development of complex memory systems. Quarq has achieved a top ranking on the LongMemEval benchmark, demonstrating the

TimeSage-MT: A Multi-Turn Benchmark for Evaluating Agentic Time Series Reasoning

Model ReleasesDGX agent

arXiv:2606.01498v1 Announce Type: cross Abstract: Time series data inform critical decisions across many real-world domains. While large language model (LLM) agents can analyze data through natural la

Unified Context Evolution for LLM Agents

AgentsDGX agent

arXiv:2606.02304v1 Announce Type: new Abstract: LLM-based agents can solve multi-step interactive tasks by combining reasoning with environment feedback, yet each episode starts from the same fixed co

Workday introduces new capabilities for building and verifying AI agents

AgentsDGX agent

Workday Inc. today announced new capabilities aimed at providing developers new ways to build on top of its platform using their own tools. During DevCon 2026, the company’s annual developer conferenc

1 Jun 2026

all agents in the future are going to need to write and execute code LangSmith Sandboxes are GA - try them out today https://docs.langchain.…

AgentsDGX agent

all agents in the future are going to need to write and execute code LangSmith Sandboxes are GA - try them out today https://docs.langchain.com/langsmith/sandboxes .@MukilLoganathan’s Interrupt keynot

Demo2: Multimodal Interactive Hybrid Agent

AgentsDGX agent

Demo2 showcases Qwen's multimodal interactive hybrid agent capabilities, likely demonstrating the integration of multiple data types (text, image, audio, video) with interactive features and hybrid pr

Demo3:Browser Agent

AgentsDGX agent

Demo3 showcases Qwen's browser agent capabilities, likely demonstrating an AI system's ability to autonomously interact with web browsers to perform tasks such as navigation, form filling, or informat

DiTTo: Scalable Order-aware All-in-One Image Restoration Agent

SafetyDGX agent

arXiv:2605.30915v1 Announce Type: new Abstract: Real-world images rarely suffer from a single degradation, and the order in which degradations are removed substantially affects the final restoration q

Fleet computer use is now available in LangSmith's APAC instance! You can now give your Fleet agents access to a virtual computer if you're …

AgentsDGX agent

LangSmith's APAC instance now supports Fleet computer use, enabling Fleet agents to access virtual computers for enhanced functionality. This feature allows users in the Asia-Pacific region to leverag

From Prompt Injection to Persistent Control: Defending Agentic Harness Against Trojan Backdoors

Model ReleasesDGX agent

arXiv:2605.31042v1 Announce Type: cross Abstract: LLM agents are evolving from conversational chatbots to operational tools in real-world workspaces. In local agentic harnesses, an LLM can read and wr

In @latentspacepod podcast, I shared my view on video generation, world models, LLMs, agents, continual learning and where the next frontier…

AgentsDGX agent

In @latentspacepod podcast, I shared my view on video generation, world models, LLMs, agents, continual learning and where the next frontier is. 1. Video models get most of their intelligence from lan

Industrializing Prediction-Powered Inference: The GLIDE Library for Reliable GenAI and Agentic Systems Evaluation

AgentsDGX agent

arXiv:2605.31278v1 Announce Type: new Abstract: Reliable evaluation of agentic systems requires unbiased estimates with valid uncertainty, but standard practice navigates between costly human annotati

← Previous
1…6970717273…299
Next →