AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,951 results
29 May 2026

Compass: Navigating Global Marine Lead Data Integration through Expert-Guided LLM Agent

AgentsDGX agent

arXiv:2605.29966v1 Announce Type: new Abstract: Marine lead (Pb) and its isotopes are critical tracers for ocean circulation and anthropogenic pollution, yet in-situ observations remain costly and spa

Graph-Enhanced Policy Optimization in LLM Agent Training

SafetyDGX agent

arXiv:2510.26270v2 Announce Type: replace Abstract: Multi-step LLM agents in interactive environments represent a crucial step toward long-horizon decision-making. To train such agents, group-based re

grok-build-0.1 is now available via the xAI API in public beta. This is the same model that powers the Grok Build CLI and excels at agentic …

AgentsDGX agent

grok-build-0.1 is now available via the xAI API in public beta. This is the same model that powers the Grok Build CLI and excels at agentic coding. Priced at 1/m input and 2/m output, it’s extremely c

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Honest Lying: Understanding Memory Confabulation in Reflexive Agents

ResearchDGX agent

arXiv:2605.29463v1 Announce Type: cross Abstract: Reflexion-style agents rely on self-generated reflections as memory, implicitly assuming that agents can accurately diagnose their own failures.We sho

i had to pack away my coding agents on monday cuz I knew I’d stay up too late if I played w activegraph during the week (yay, it’s Friday!) …

AgentsDGX agent

i had to pack away my coding agents on monday cuz I knew I’d stay up too late if I played w activegraph during the week (yay, it’s Friday!) gautham kept playing with it and just showed me a custom UI

I literally haven’t typed anything in weeks Since I started using Grok’s speech-to-text in Hermes Agent... I just talk It transcribes everyt…

AgentsDGX agent

I literally haven’t typed anything in weeks Since I started using Grok’s speech-to-text in Hermes Agent... I just talk It transcribes everything perfectly Every word. Every single time My thoughts flo

I've been using state-of-the-art models to teach small models running on my computer how I work. The result : a personal agent that runs my …

AgentsDGX agent

I've been using state-of-the-art models to teach small models running on my computer how I work. The result : a personal agent that runs my inbox, my deal pipeline, my blog, my calendar, & my research

Learning to Choose: An Empowerment-Guided Multi-Agent System with semantic communication for Adaptive Method Selection

SafetyDGX agent

arXiv:2605.30042v1 Announce Type: new Abstract: Automating scientific computing workflows requires more than generating executable code: autonomous systems must also select appropriate computational s

Most people training agentic LLMs with RL right now have a silently broken training loop and have no idea. Here's the trap: single-turn RL w…

AgentsDGX agent

Most people training agentic LLMs with RL right now have a silently broken training loop and have no idea. Here's the trap: single-turn RL works beautifully. Clean curves, sane rewards, everything con

Open Source Browser Agent That Learns and Repeats Workflows

AgentsDGX agent

An open-source browser with built-in AI agents that emphasizes privacy and automation, enabling task automation through natural language without coding. The browser supports multiple AI providers incl

PersonaAgent: Bridging Memory and Action for Personalized LLM Agents

SafetyDGX agent

arXiv:2506.06254v2 Announce Type: replace Abstract: Large Language Model (LLM) empowered agents have recently emerged as advanced paradigms that exhibit impressive capabilities in a wide range of doma

Sources: Microsoft is working on an app that will include GitHub Copilot, Copilot chat, Copilot Cowork, and a new agentic workflow tool called Autopilot (Sebastian Herrera/Fortune)

AgentsDGX agent

Sebastian Herrera / Fortune: Sources: Microsoft is working on an app that will include GitHub Copilot, Copilot chat, Copilot Cowork, and a new agentic workflow tool called Autopilot — Microsoft needs

STAMP: Training Explicit Memory for Mobile GUI Agents in Controllable and Scalable Virtual Environments

Model ReleasesDGX agent

arXiv:2605.29324v1 Announce Type: new Abstract: Mobile GUI agents excel at immediate reactive control but frequently fail in realistic, long-horizon tasks that require memory. This failure stems from

Worked on some code this morning using Opus 4.8 and so far I'm really liking it. Much more cooperative than 4.7 and less 'over agentic'. Sto…

AgentsDGX agent

Worked on some code this morning using Opus 4.8 and so far I'm really liking it. Much more cooperative than 4.7 and less 'over agentic'. Stops and asks for my input when needed in places 4.7 (and GPT

28 May 2026

Asana acquires StackAI, a no-code platform for building AI agents, for 75M as part of Asana's broader AI pivot; PitchBook: StackAI raised ~20M (Russell Brandom/TechCrunch)

AgentsDGX agent

Russell Brandom / TechCrunch: Asana acquires StackAI, a no-code platform for building AI agents, for 75M as part of Asana's broader AI pivot; PitchBook: StackAI raised ~20M — Asana has acquired the wo

Coherence Collapse: Diagnosing Why Code Agents Fail After Reaching the Right Code

Model ReleasesDGX agent

arXiv:2603.24631v2 Announce Type: replace-cross Abstract: Code agents resolve 65-70% of SWE-bench Verified issues, but Pass@1 cannot tell us why the rest fail, and, as we show, capable-model failures

Deploying a Hermes Agent with Fly, Modal, OpenRouter, & Cloudflare 02:43 Managed vs VPS 06:08 Architecture 15:14 Setup 22:23 Deployment 28:0…

AgentsDGX agent

Deploying a Hermes Agent with Fly, Modal, OpenRouter, & Cloudflare 02:43 Managed vs VPS 06:08 Architecture 15:14 Setup 22:23 Deployment 28:00 Access / OIDC 39:57 Hermes and Open WebUI 52:07 Cloudflare

Discovery Agents for Real-Time Analytics: Toward Proactive Insight Systems

AgentsDGX agent

arXiv:2605.27571v1 Announce Type: new Abstract: Modern analytics systems are fundamentally reactive, requiring users to define queries over increasingly complex and continuously evolving data. In real

Hierarchical Prompt-Domain Control and Learning for Resource-Constrained Agentic Language Models

AgentsDGX agent

arXiv:2605.27703v1 Announce Type: new Abstract: Large Language Models are increasingly deployed inside agentic systems, where they must follow structured protocols, adapt to evolving states, and opera

How we built Cloudflare's data platform and an AI agent on top of it

AgentsDGX agent

Cloudflare describes the architecture and development of their unified data platform that integrates data from across their global network, along with an AI agent built on top to enable intelligent qu

Okta reports Q1 revenue up 11% YoY to 765M, vs. 752M est., says the agentic AI build-out is spiking demand for its identity tools; OKTA jumps 8%+ after hours (Samantha Subin/CNBC)

AgentsDGX agent

Samantha Subin / CNBC: Okta reports Q1 revenue up 11% YoY to 765M, vs. 752M est., says the agentic AI build-out is spiking demand for its identity tools; OKTA jumps 8%+ after hours — Okta beat Wall St

President and Head of AI at Replit, @pirroh is the architect of Replit Agent and former Head of Applied Research at Google X. See him take t…

AgentsDGX agent

President and Head of AI at Replit, @pirroh is the architect of Replit Agent and former Head of Applied Research at Google X. See him take the stage with @refikanadol on day two of Vibecon. NYC, June

📢Qwen3.7-Max just hit #3 on ITbench-AA — a fresh benchmark testing how well models handle real-world enterprise IT tasks, agentic-style. 🔧…

Model ReleasesDGX agent

📢Qwen3.7-Max just hit #3 on ITbench-AA — a fresh benchmark testing how well models handle real-world enterprise IT tasks, agentic-style. 🔧Agentic era, go with Qwen.🏃🏃 Artificial Analysis and IBM Resea

Roles with Rails: Contract-Preserving Role Evolution in Multi-Agent Structured Reasoning

AgentsDGX agent

arXiv:2605.28433v1 Announce Type: new Abstract: Role-based LLM multi-agent systems need adaptive role pools, yet adapting such systems is not merely a matter of prompt optimization: roles often carry

Structured Agent Distillation for Large Language Model

SafetyDGX agent

arXiv:2505.13820v5 Announce Type: replace-cross Abstract: Large language models (LLMs) exhibit strong capabilities as decision-making agents by interleaving reasoning and actions, as seen in ReAct-sty

VeriTrip: A Verifiable Benchmark for Travel Planning Agents over Unstructured Web Corpora

Model ReleasesDGX agent

arXiv:2605.28683v1 Announce Type: new Abstract: Existing benchmarks have laid the foundation for travel planning agents by establishing API-centric paradigms. However, as the capabilities of Autonomou

27 May 2026

BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

Model ReleasesDGX agent

arXiv:2603.03194v2 Announce Type: replace Abstract: Current code-agent benchmarks primarily evaluate localized issue resolution within a single target repository, leaving under-tested many software en

Building self-improving tax agents with Codex

AgentsDGX agent

This article describes how OpenAI's Codex model can be used to build autonomous tax agents capable of self-improvement through code generation and execution. The work demonstrates using large language

Communication Gain and Delay Cost Under Cross-Timestep Delays in Cooperative Multi-Agent Reinforcement Learning

AgentsDGX agent

arXiv:2604.03785v2 Announce Type: replace Abstract: Communication is essential for coordination in cooperative multi-agent reinforcement learning under partial observability, yet cross-timestep delays

Counterfactual Credit Policy Optimization for Multi-Agent Collaboration

SafetyDGX agent

arXiv:2603.21563v2 Announce Type: replace Abstract: Collaborative multi-agent large language models (LLMs) can solve complex reasoning tasks by decomposing roles, but reinforcement learning for such s

Demis Hassabis says he still broadly expects AGI around 2030, though he now sees 2029 as a possibility, and 2026's 'agentic era' is a 'bit like a practice run' (Ina Fried/Axios)

AgentsDGX agent

Ina Fried / Axios: Demis Hassabis says he still broadly expects AGI around 2030, though he now sees 2029 as a possibility, and 2026's “agentic era” is a “bit like a practice run” — Google DeepMind CEO

From data overload to actionable insights: How Verizon Connect scaled agentic AI to 100,000 users

AgentsDGX agent

In this post, we show you how Verizon Connect built and scaled an agentic AI solution to transform overwhelming fleet data into clear, actionable insights for 100,000 users daily. We walk you through

I really appreciate the lessons and technical ideas @samaysham & team were able to share about their tax agent system, which learns from pro…

AgentsDGX agent

I really appreciate the lessons and technical ideas @samaysham & team were able to share about their tax agent system, which learns from production traces to self-improve via detailed tracing tightly

Interactive Agents: Simulating Counselor-Client Psychological Counseling via Role-Playing LLM-to-LLM Interactions

AgentsDGX agent

arXiv:2408.15787v2 Announce Type: replace Abstract: Creating effective dialogue systems for mental health support requires high-quality multi-turn counseling dialogue data, yet collecting real counsel

JobBench: Aligning Agent Work With Human Will

Model ReleasesDGX agent

arXiv:2605.26329v1 Announce Type: new Abstract: Current benchmarks for occupational AI agents are scoped primarily by economic values, telling a replacement story. We introduce JobBench, which evaluat

RICE-PO: Turning Retrieval Interactions into Credit Signals for Reasoning Agents

SafetyDGX agent

arXiv:2605.26352v1 Announce Type: new Abstract: Retrieval is increasingly moving from one-shot matching toward interactive reasoning, where language agents iteratively inspect evidence, reformulate qu

SPEAR: Code-Augmented Agentic Prompt Optimization

AgentsDGX agent

arXiv:2605.26275v1 Announce Type: new Abstract: Automatic prompt engineering (APE) rewrites prompts to improve downstream task performance, but existing APE loops treat the optimizer itself as a fixed

SWE-Adept: An LLM-Based Agentic Framework for Deep Codebase Analysis and Structured Issue Resolution

Model ReleasesDGX agent

arXiv:2603.01327v2 Announce Type: replace-cross Abstract: Large language models (LLMs) exhibit strong performance on self-contained programming tasks. However, they still struggle with repository-leve

Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction

SafetyDGX agent

arXiv:2505.11063v3 Announce Type: replace Abstract: LLM-based agents solve complex tasks through iterative reasoning, tool use, and environment interaction, where each intermediate thought directly sh

Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization

Model ReleasesDGX agent

arXiv:2605.26457v1 Announce Type: cross Abstract: AI coding agents are increasingly used to write real-world software, but ensuring that their outputs are correct remains a fundamental challenge. Form

26 May 2026

Automated Detection and Classification of Delusion-related Content in Naturalistic Audio Diaries Using Multi-Agent Language Models

AgentsDGX agent

arXiv:2605.24755v1 Announce Type: new Abstract: Speech monologues recorded in naturalistic settings provide opportunities to characterize mental illness phenomenology and detect symptom exacerbation.

Board Mode is here. Instead of one long agent session, you now get a creative workspace where research, slides, websites, apps, and design d…

AgentsDGX agent

Board Mode is here. Instead of one long agent session, you now get a creative workspace where research, slides, websites, apps, and design drafts can live side by side. Try it now: https://agent.ii.in

Inference-Time Backdoors via Chat Templates: From LLM Supply Chains to Agentic System Compromise

AgentsDGX agent

arXiv:2602.04653v4 Announce Type: replace-cross Abstract: Open-weight language models are increasingly used in production settings, raising new security challenges. One prominent threat is backdoor at

Just built an insane new agent skill. It can perfectly extract slides from YT videos, then write notes, images, transcripts, and slides into…

AgentsDGX agent

Just built an insane new agent skill. It can perfectly extract slides from YT videos, then write notes, images, transcripts, and slides into Obsidian vaults. An HTML artifact allows me to navigate and

Methods for Formal Verification of Agent Skills: Three Layers Toward a Mechanically Checkable Capability-Containment Proof

AgentsDGX agent

arXiv:2605.23951v1 Announce Type: new Abstract: The companion paper introduced a four-level verification lattice on agent-skill manifests (unverified, declared, tested, formal) and left the top level

Novee debuts Agentic Fix, pushing pentest findings into Claude, Copilot and Cursor

Model ReleasesDGX agent

Artificial intelligence penetration testing startup Novee Cyber Security Ltd. today launched Agentic Fix, a new capability that pushes validated exploit findings directly into the AI coding agents dev

VineLM: Trie-Based Fine-Grained Control for Agentic Workflows

AgentsDGX agent

arXiv:2605.23914v1 Announce Type: cross Abstract: Agentic workflows interleave configurable LLM stages with tool stages and often include retries or refinement loops. Existing workflow managers profil

Why Agentic Theorem Prover Works: A Statistical Provability Theory of Mathematical Reasoning Models

AgentsDGX agent

arXiv:2602.10538v3 Announce Type: replace-cross Abstract: Agentic theorem provers combine a reasoning model, retrieval, search, and a proof assistant verifier, yet it remains unclear which components

Why Your Deep Research Agent Fails? On Hallucination Evaluation in Full Research Trajectory

AgentsDGX agent

arXiv:2601.22984v2 Announce Type: replace Abstract: Diagnosing failure patterns in Deep Research Agents (DRAs) remains a critical challenge. Existing benchmarks predominantly rely on end-to-end evalua

25 May 2026

A Proactive Multi-Agent Dialogue Framework for Assessing Social Language Disorder Traits in Autism

AgentsDGX agent

arXiv:2605.22993v1 Announce Type: cross Abstract: Characteristic linguistic behaviors associated with Social Language Disorder (SLD) in autism spectrum disorder, including echoic repetition, pronoun d

(I'm firmly on team red/green TDD for agent code, I like having a test suite that protects against them breaking old features when they make…

AgentsDGX agent

(I'm firmly on team red/green TDD for agent code, I like having a test suite that protects against them breaking old features when they make new changes - https://simonwillison.net/guides/agentic-engi

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents

Model ReleasesDGX agent

arXiv:2605.23657v1 Announce Type: new Abstract: Skills, i.e., structured workflow instructions distilled for large language models (LLMs), are becoming an increasingly important mechanism for improvin

PACE: Two-Timescale Self-Evolution for Small Language Model Agents

Model ReleasesDGX agent

arXiv:2605.23019v1 Announce Type: new Abstract: Deploying language-model agents in production often requires substantial compute and human effort to tune prompts, parsers, validators, and other compon

PhotoFlow: Agentic 3D Virtual Photography Missions

Model ReleasesDGX agent

arXiv:2605.23771v1 Announce Type: cross Abstract: Virtual photography asks an agent to enter a prepared 3D scene with no preselected camera pose or reference image, infer a suitable shot from scene in

SciNet: Evaluating AI Agents in Relation-Aware Scientific Literature Retrieval

ResearchDGX agent

arXiv:2601.03260v2 Announce Type: replace-cross Abstract: AI agents have seen widespread adoption in information retrieval for scientific research, giving rise to tools such as Deep Research. However,

23 May 2026

last night i got an agent to fork itself, propose a modification to itself on the fork, run through tests (sandbox, etc), and only accept th…

AgentsDGX agent

Yohei Nakajima describes a technical demonstration where an AI agent was configured to create a fork of itself, propose modifications to the forked version, execute tests in a sandboxed environment, a

Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents

Model ReleasesDGX agent

arXiv:2605.21768v1 Announce Type: new Abstract: Memory-augmented LLM agents enable interactions that extend beyond finite context windows by storing, updating, and reusing information across sessions.

22 May 2026

AMD CEO Lisa Su projects the CPU market will grow over 35% annually through 2031, up from 3% to 4% historically, driven by AI inference and agentic AI demand (Cheng Ting-Fang/Nikkei Asia)

AgentsDGX agent

Cheng Ting-Fang / Nikkei Asia: AMD CEO Lisa Su projects the CPU market will grow over 35% annually through 2031, up from 3% to 4% historically, driven by AI inference and agentic AI demand — TAIPEI —

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI

Model ReleasesDGX agent

arXiv:2603.14987v2 Announce Type: replace Abstract: Agentic AI systems increasingly act through tool-augmented, multi-step workflows whose failures (unsafe tool use, unauthorised actions, social harm)

Declarative Data Services: Structured Agentic Discovery for Composing Data Systems

Model ReleasesDGX agent

arXiv:2605.20690v1 Announce Type: new Abstract: Agentic discovery has shown that LLM-driven search can find novel algorithms, designs, and code under benchmark conditions. Translating the paradigm to

← Previous
1…8384858687…300
Next →