AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,964 results
1 Jul 2026

AxDafny: Agentic Verified Code Generation in Dafny

Model ReleasesDGX agent

arXiv:2606.32007v1 Announce Type: new Abstract: We study agentic code generation in Dafny, where a model must generate both executable code and the proof artifacts for verification. We present AxDafny

do you know what you pay for in agentic workloads? cached tokens! session with 50+ tool calls -> prompt is billed 50 times all providers giv…

Model ReleasesDGX agent

do you know what you pay for in agentic workloads? cached tokens! session with 50+ tool calls -> prompt is billed 50 times all providers give 1/5 cached discount for GLM-5.2 we at @FireworksAI_HQ drop

ECHO: Prune to act, trace to learn with selective turn memory in agentic RL

SafetyDGX agent

arXiv:2606.31650v1 Announce Type: cross Abstract: Long-horizon language agents must repeatedly interact with tools, accumulate evidence, and make decisions under bounded context windows. Existing cont

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

good post on the value of knowing whats going on in your harness 'For agents within an actual product, I have gravitated to using LangChain …

Model ReleasesDGX agent

good post on the value of knowing whats going on in your harness 'For agents within an actual product, I have gravitated to using LangChain DeepAgents' The frontier model lab harnesses are amazing - d

IterCAD: An Iterative Multimodal Agent for Visually-Grounded CAD Generation and Editing

SafetyDGX agent

arXiv:2606.13368v2 Announce Type: replace Abstract: Computer-Aided Design is pivotal in modern manufacturing, yet existing automated methods predominantly rely on open-loop, one-shot generation, creat

LabGuard: Grounding Natural-Language Laboratory Rules into Runtime Guards for Embodied Laboratory Agents

SafetyDGX agent

arXiv:2606.31045v1 Announce Type: new Abstract: Scientific embodied agents are increasingly capable of carrying out laboratory procedures, but executing these procedures safely in dynamic laboratory e

MECoBench: A Systematic Study of Multimodal Agent Collaboration in Embodied Environments

Model ReleasesDGX agent

arXiv:2606.31966v1 Announce Type: cross Abstract: Recent multimodal large language models (MLLMs) have strong potential as embodied agents, but their ability to collaborate in visually grounded enviro

Our research team has 9 papers at ICML next week! Spanning the full stack from frontier agents to GPU kernels, we're excited to share what o…

HardwareDGX agent

Our research team has 9 papers at ICML next week! Spanning the full stack from frontier agents to GPU kernels, we're excited to share what our researchers and collaborators have been working on. If yo

PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks

Model ReleasesDGX agent

arXiv:2606.31154v1 Announce Type: cross Abstract: Creating and editing slides is a rich, multimodal activity that is ubiquitous in professional and educational settings, making it an ideal testbed for

QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents

SafetyDGX agent

arXiv:2606.32034v1 Announce Type: cross Abstract: LLM agents increasingly act over long horizons, where a single trajectory can contain hundreds or thousands of actions. In these settings, outcome-onl

Smart charging of large fleets of Electric Vehicles: Independent Multi-Agent Reinforcement Learning approaches

SafetyDGX agent

arXiv:2606.31347v1 Announce Type: new Abstract: The electrification of transportation through electric vehicles introduces new challenges for power grid management, such as increased peak demand, volt

TreeAgent: A Generalizable Multi-Agent Framework for Automated Bias Labeling in Forestry via Compiled Expert Rules and Vision-Language Models

SafetyDGX agent

arXiv:2606.31976v1 Announce Type: new Abstract: Human-labeled data are widely used as reference annotations in ML, despite known variability across annotators in many expert-driven domains. In additio

TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2606.32017v1 Announce Type: cross Abstract: Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edits, navigation commands, and objec

30 Jun 2026

Budgeted Act-or-Defer Multi-Agent LLM Deliberation with Local Reliability Bounds

SafetyDGX agent

arXiv:2606.29654v1 Announce Type: new Abstract: Multi-agent deliberation among LLMs can improve reasoning, but deployment requires deciding when the current answer is reliable enough to act on and whe

Building to the Test: Coding Agents Deliver What You Check, Not What You Requested

Model ReleasesDGX agent

arXiv:2606.28430v1 Announce Type: cross Abstract: Benchmarks are widely used to evaluate task completion by Large Language Models (LLMs), but this approach has accumulated construction-validity proble

Forensic Trajectory Signatures for Agent Memory Poisoning Detection

Model ReleasesDGX agent

arXiv:2606.30566v1 Announce Type: cross Abstract: We discover a behavioral invariant in LLM agents under persistent memory poisoning: in architectures where routing information is retrieved through ob

From Tool Connection to Execution Control: Benchmarking Security Invariants in MCP-Style Agent Runtimes

Model ReleasesDGX agent

arXiv:2606.29073v1 Announce Type: cross Abstract: Model Context Protocol (MCP)-style ecosystems give language-model applications a practical connection layer for tools, resources, prompts, and transpo

HiComm: Hierarchical Communication for Multi-agent Reinforcement Learning

SafetyDGX agent

arXiv:2606.29126v1 Announce Type: new Abstract: Cooperative multi-agent reinforcement learning (MARL) often relies on communication to mitigate partial observability, yet most existing protocols treat

Hotels, tour operators, and travel agencies rush to launch proprietary online tools and loyalty schemes to fend off future competition from AI travel agents (Stephanie Stacey/Financial Times)

IndustryDGX agent

Stephanie Stacey / Financial Times: Hotels, tour operators, and travel agencies rush to launch proprietary online tools and loyalty schemes to fend off future competition from AI travel agents — Chatb

Looking Is Not Picking: An Attention-Segment Account of Tool-Selection Failures in LLM Agents

Local AiDGX agent

arXiv:2606.16364v2 Announce Type: replace Abstract: LLM agents mis-call tools, and the natural guess is that the model failed to see the right tool in a crowded harness. We show the opposite through a

LUMEN: Cost-Transparent Multi-Agent Pipeline for Automated Systematic Review and Meta-Analysis

Model ReleasesDGX agent

arXiv:2606.28362v1 Announce Type: cross Abstract: Systematic reviews and meta-analyses (SR/MA) remain the gold standard for evidence synthesis, yet completing one typically requires 67 weeks and subst

MemLeak: Diagnosing Information Leaks in Multimodal Agent Memory

Model ReleasesDGX agent

arXiv:2606.29788v1 Announce Type: new Abstract: When a multimodal AI agent is asked to forget a fact, current memory systems usually delete the text entry and report success. We find that the fact can

meta-pipe: An LLM-agent pipeline for end-to-end automated systematic review and meta-analysis

Model ReleasesDGX agent

arXiv:2606.28363v1 Announce Type: cross Abstract: Objective: To describe the architecture and design rationale of meta-pipe, an open-source large language model (LLM)-agent pipeline that integrates th

Metric Aggregation Divergence: A Hidden Validity Threat in Agent-Based Policy Optimization and a Contractual Remedy

SafetyDGX agent

arXiv:2606.29038v1 Announce Type: cross Abstract: Metric aggregation divergence (MAD) is the silent inconsistency that arises when distinct pipeline stages in an agent-based model coupled with a multi

Multi-Agent Route Planning as a QUBO Problem

Model ReleasesDGX agent

arXiv:2602.07913v2 Announce Type: replace Abstract: Multi-Agent Route Planning considers selecting vehicles, each associated with a single predefined route, such that route-level coverage utility is m

OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks

Model ReleasesDGX agent

arXiv:2606.29537v1 Announce Type: new Abstract: Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world computer use, limiting their ability to

Our Founder and Chief Scientist @EdoLiberty just kicked off the Search & Retrieval track @aiDotEngineer. 'Agents don't need to be smarter. T…

ToolsDGX agent

Our Founder and Chief Scientist @EdoLiberty just kicked off the Search & Retrieval track @aiDotEngineer. 'Agents don't need to be smarter. They need a Knowledge Layer.' We've got a big update on this

ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration

ApplicationsDGX agent

ScarfBench is a benchmarking framework designed to evaluate AI agents' capabilities in migrating enterprise Java applications to modern frameworks. The benchmark likely assesses how well AI systems ca

We’re introducing GeneBench-Pro, a research-level benchmark for a harder kind of AI progress: how well agents can navigate messy biological …

Model ReleasesDGX agent

We’re introducing GeneBench-Pro, a research-level benchmark for a harder kind of AI progress: how well agents can navigate messy biological data, choose the right analysis path, and make judgment call

When Does Overlap Help? OSU-Mem and a Cell-Conditional Analysis of Trajectory Memory for LLM Agents

Model ReleasesDGX agent

arXiv:2606.28376v1 Announce Type: cross Abstract: Long-horizon large language model (LLM) agents accumulate interaction trajectories that quickly exceed any practical prompt budget, and existing memor

29 Jun 2026

ATOD: Annealed Turn-aware On-policy Distillation for Multi-turn Autonomous Agents

SafetyDGX agent

arXiv:2606.27814v1 Announce Type: new Abstract: Training small language-model agents for long-horizon interactive tasks requires both fast imitation and reward-driven improvement. On-policy distillati

Google Cloud Japan ブログにインタビュー記事を掲載いただきました。 「Sakana Fugu」のサービス基盤として、GoogleのEnterprise Agent Platformを全面採用した経緯について、プロジェクトを率いたチーフサイエンティストのYujin…

Model ReleasesDGX agent

Google Cloud Japan ブログにインタビュー記事を掲載いただきました。 「Sakana Fugu」のサービス基盤として、GoogleのEnterprise Agent Platformを全面採用した経緯について、プロジェクトを率いたチーフサイエンティストのYujin Tangをはじめ、3名のサイエンティストが開発の舞台裏を語っています。 「Sakana Fugu はスポーツに例えるな

Multimodal Evaluator Preference Collapse: Cross-Modal Coupling in Self-Evolving Agents

Model ReleasesDGX agent

arXiv:2606.16682v3 Announce Type: replace-cross Abstract: When AI agents use language models to evaluate their own outputs in a feedback loop, systematic biases emerge. We show that Evaluator Preferen

QuantV2X: A Fully Quantized Multi-Agent System for Cooperative Perception

Model ReleasesDGX agent

arXiv:2509.03704v2 Announce Type: replace Abstract: Cooperative perception through Vehicle-to-Everything (V2X) communication offers significant potential for enhancing vehicle perception by mitigating

this has to be because coding agents change the engineering math on how it is to work with or port a legacy codebase, right? anyone at Riot …

ToolsDGX agent

this has to be because coding agents change the engineering math on how it is to work with or port a legacy codebase, right? anyone at Riot able to confirm? League of Legends Classic is coming as a ne

Voice agents, now on Vercel. Realtime, speech and transcription are now live on AI Gateway. Build with 𝚞𝚜𝚎𝚁𝚎𝚊𝚕𝚝𝚒𝚖𝚎, 𝚐𝚎𝚗𝚎𝚛𝚊…

ToolsDGX agent

Vercel has launched voice agent capabilities on its AI Gateway platform, enabling developers to build real-time applications with speech and transcription features. The announcement indicates new inte

27 Jun 2026

I put together a new article on setting up local coding agents with open-weight models. Everything runs 100% locally. I thought it might be …

Model ReleasesDGX agent

I put together a new article on setting up local coding agents with open-weight models. Everything runs 100% locally. I thought it might be useful putting this together because many people asked me ab

26 Jun 2026

Diagnosing Task Insensitivity in Language Agents

SafetyDGX agent

arXiv:2606.26918v1 Announce Type: new Abstract: Large language models can serve as capable long-horizon agents, but their out-of-distribution (OOD) generalization remains weak. We identify a key sourc

EGG: An Expert-Guided Agent Framework for Kernel Generation

HardwareDGX agent

arXiv:2606.26758v1 Announce Type: new Abstract: High-performance GPU kernels are critical for reducing the exponentially growing computational costs of large language models (LLMs), but their developm

EVOM: Agentic Meta-Evolution of Actor-Critic Architectures for Reinforcement Learning

SafetyDGX agent

arXiv:2606.26327v1 Announce Type: cross Abstract: In actor-critic reinforcement learning, network architectures are typically manually designed. Automating this design is challenging because each cand

Improving General Role-Playing Agents via Psychology-Grounded Reasoning and Role-Aware Policy Optimization

SafetyDGX agent

arXiv:2606.27025v1 Announce Type: new Abstract: Building general-purpose role-playing agents that faithfully portray any character from a natural-language profile remains challenging. The dominant par

Memory Depth, Not Memory Access: Selective Parametric Consolidation for Long-Running Language Agents

Model ReleasesDGX agent

arXiv:2606.26806v1 Announce Type: new Abstract: Long-running language agents need more than memory access. Retrieval systems can fetch past facts at query time, but they do not decide which experience

Post-Training Recipe, More Than Model Family, Shapes Multi-Agent LLM Conversational Behavior

Model ReleasesDGX agent

arXiv:2606.20632v2 Announce Type: replace-cross Abstract: Multi-LLM systems use multiple language models to deliberate, judge each other's outputs, or coordinate as agents. Their value depends on the

R2D-RL: A RoboCup 2D Soccer Environment for Multi-Agent Reinforcement Learning

Model ReleasesDGX agent

arXiv:2606.18786v2 Announce Type: replace Abstract: Robot soccer is a challenging testbed for multi-agent reinforcement learning because it combines partial observability, cooperative and adversarial

Semantic Early-Stopping for Iterative LLM Agent Loops

SafetyDGX agent

arXiv:2606.27009v1 Announce Type: new Abstract: Multi-agent large language model (LLM) loops, for example a Writer that drafts and a Critic that revises, are almost always terminated by a fixed iterat

The strongest models are gated and access is granted only to a select few. Hermes Agent now exposes MoA presets as virtual models, giving yo…

Model ReleasesDGX agent

The strongest models are gated and access is granted only to a select few. Hermes Agent now exposes MoA presets as virtual models, giving you capabilities beyond the publicly available frontier: 8% hi

When Actions Go Off-Task: Detecting and Correcting Misaligned Actions in Computer-Use Agents

Model ReleasesDGX agent

arXiv:2602.08995v2 Announce Type: replace Abstract: Computer-use agents (CUAs) have made tremendous progress in the past year, yet they still frequently produce misaligned actions that deviate from th

When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models

SafetyDGX agent

arXiv:2606.27288v1 Announce Type: new Abstract: Multi-model LLM systems such as routing, voting, cascades, fusion, and mixture-of-agents are used to beat single-model accuracy. We show that their gain

Where Do CoT Training Gains Land in LLM based Agents?

ResearchDGX agent

arXiv:2606.26935v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning is widely used in language-model agents, but prior work has shown that verbalized CoT is not always faithful and may in

25 Jun 2026

Adaptive-Horizon Conflict-Based Search for Closed-Loop Multi-Agent Path Finding

SafetyDGX agent

arXiv:2602.12024v3 Announce Type: replace Abstract: Multi-Agent Path Finding (MAPF) is a core coordination problem for large robot fleets in automated warehouses and logistics. Existing approaches are

ASAP: Agent-System Co-Design for Wall-Clock-Centered Auto HPO Research for ML Experiments

SafetyDGX agent

arXiv:2606.25207v1 Announce Type: cross Abstract: Hyperparameter Optimization (HPO) is essential for maximizing machine learning model performance, and its core challenge is sample efficiency: finding

General Intuition, which trains AI agents in spatial reasoning via gameplay footage, raised 320M led by Khosla at a 2.3B valuation, for $454M in total funding (Rebecca Bellan/TechCrunch)

ApplicationsDGX agent

Rebecca Bellan / TechCrunch: General Intuition, which trains AI agents in spatial reasoning via gameplay footage, raised 320M led by Khosla at a 2.3B valuation, for $454M in total funding — As soon as

Notion plans to shut down its Gmail client Notion Mail on September 22 and go 'all in' on AI agents to run inboxes, saying 50%+ of users do not open the inbox (Zac Hall/9to5Mac)

IndustryDGX agent

Zac Hall / 9to5Mac: Notion plans to shut down its Gmail client Notion Mail on September 22 and go “all in” on AI agents to run inboxes, saying 50%+ of users do not open the inbox — Last year, Notion e

Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents

SafetyDGX agent

arXiv:2606.25852v1 Announce Type: new Abstract: Group-based reinforcement learning effectively post-trains LLM agents for long-horizon, sparse-reward tasks by deriving step-level credit from trajector

24 Jun 2026

3/3 We took 7 weak agents (ranks 7-13, none scoring >45) from the leaderboard & merged them into 1 report/task. Essentially boosting for dee…

Model ReleasesDGX agent

3/3 We took 7 weak agents (ranks 7-13, none scoring >45) from the leaderboard & merged them into 1 report/task. Essentially boosting for deep research. The result: New #1 DRB II TotalScore of 64.38. F

At Hugging Face we've been building our own agent that we use via Slack (Moon Bot). Honestly, building your own is quite simple and you'll b…

Model ReleasesDGX agent

At Hugging Face we've been building our own agent that we use via Slack (Moon Bot). Honestly, building your own is quite simple and you'll be happy you did: any model you want (self-hosted if needed),

Can Language Model Agents be Helpful Circuit Explainers in Mechanistic Interpretability?

Model ReleasesDGX agent

arXiv:2606.24026v1 Announce Type: new Abstract: Mechanistic interpretability has made substantial progress in automatically localizing circuits, but explaining what localized components do remains lab

Customer engagement service MoEngage acquires Aampe, whose AI agents help brands personalize messaging, for 'tens of millions', per a source; Aampe raised ~$28M (Jagmeet Singh/TechCrunch)

IndustryDGX agent

Jagmeet Singh / TechCrunch: Customer engagement service MoEngage acquires Aampe, whose AI agents help brands personalize messaging, for “tens of millions”, per a source; Aampe raised ~$28M — Indian cu

DeepBD: A Grounded Agentic Workflow for Variant Prioritization and Diagnosis of Genetic Birth Defects

Model ReleasesDGX agent

arXiv:2606.24779v1 Announce Type: cross Abstract: Birth defects are a major cause of fetal loss, neonatal morbidity and long-term disability. In the subset with suspected genetic etiologies, exome and

London-based Isometric, which develops AI agents that work with human verifiers to automate industrial certification, raised a €34M Series A led by AVP (David Cendon Garcia/EU-Startups)

IndustryDGX agent

David Cendon Garcia / EU-Startups: London-based Isometric, which develops AI agents that work with human verifiers to automate industrial certification, raised a €34M Series A led by AVP — Isometric h

← Previous
1…128129130131132…300
Next →