AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,289 results
Safety

QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents

DGX agent

arXiv:2606.32034v1 Announce Type: cross Abstract: LLM agents increasingly act over long horizons, where a single trajectory can contain hundreds or thousands of actions. In these settings, outcome-onl

safetyarxiv-cs-ai
1 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Smart charging of large fleets of Electric Vehicles: Independent Multi-Agent Reinforcement Learning approaches

DGX agent

arXiv:2606.31347v1 Announce Type: new Abstract: The electrification of transportation through electric vehicles introduces new challenges for power grid management, such as increased peak demand, volt

safetyarxiv-cs-ai
1 Jul 2026
Safety

TreeAgent: A Generalizable Multi-Agent Framework for Automated Bias Labeling in Forestry via Compiled Expert Rules and Vision-Language Models

DGX agent

arXiv:2606.31976v1 Announce Type: new Abstract: Human-labeled data are widely used as reference annotations in ML, despite known variability across annotators in many expert-driven domains. In additio

safetyarxiv-cs-ai
1 Jul 2026
Safety

TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning

DGX agent

arXiv:2606.32017v1 Announce Type: cross Abstract: Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edits, navigation commands, and objec

safetyarxiv-cs-ai
1 Jul 2026
Safety

Budgeted Act-or-Defer Multi-Agent LLM Deliberation with Local Reliability Bounds

DGX agent

arXiv:2606.29654v1 Announce Type: new Abstract: Multi-agent deliberation among LLMs can improve reasoning, but deployment requires deciding when the current answer is reliable enough to act on and whe

safetyarxiv-cs-ai
30 Jun 2026
Model Releases

Building to the Test: Coding Agents Deliver What You Check, Not What You Requested

DGX agent

arXiv:2606.28430v1 Announce Type: cross Abstract: Benchmarks are widely used to evaluate task completion by Large Language Models (LLMs), but this approach has accumulated construction-validity proble

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Forensic Trajectory Signatures for Agent Memory Poisoning Detection

DGX agent

arXiv:2606.30566v1 Announce Type: cross Abstract: We discover a behavioral invariant in LLM agents under persistent memory poisoning: in architectures where routing information is retrieved through ob

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

From Tool Connection to Execution Control: Benchmarking Security Invariants in MCP-Style Agent Runtimes

DGX agent

arXiv:2606.29073v1 Announce Type: cross Abstract: Model Context Protocol (MCP)-style ecosystems give language-model applications a practical connection layer for tools, resources, prompts, and transpo

model-releasesarxiv-cs-ai
30 Jun 2026
Safety

HiComm: Hierarchical Communication for Multi-agent Reinforcement Learning

DGX agent

arXiv:2606.29126v1 Announce Type: new Abstract: Cooperative multi-agent reinforcement learning (MARL) often relies on communication to mitigate partial observability, yet most existing protocols treat

safetyarxiv-cs-ai
30 Jun 2026
Local Ai

Looking Is Not Picking: An Attention-Segment Account of Tool-Selection Failures in LLM Agents

DGX agent

arXiv:2606.16364v2 Announce Type: replace Abstract: LLM agents mis-call tools, and the natural guess is that the model failed to see the right tool in a crowded harness. We show the opposite through a

local-aiarxiv-cs-ai
30 Jun 2026
Model Releases

LUMEN: Cost-Transparent Multi-Agent Pipeline for Automated Systematic Review and Meta-Analysis

DGX agent

arXiv:2606.28362v1 Announce Type: cross Abstract: Systematic reviews and meta-analyses (SR/MA) remain the gold standard for evidence synthesis, yet completing one typically requires 67 weeks and subst

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

MemLeak: Diagnosing Information Leaks in Multimodal Agent Memory

DGX agent

arXiv:2606.29788v1 Announce Type: new Abstract: When a multimodal AI agent is asked to forget a fact, current memory systems usually delete the text entry and report success. We find that the fact can

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

meta-pipe: An LLM-agent pipeline for end-to-end automated systematic review and meta-analysis

DGX agent

arXiv:2606.28363v1 Announce Type: cross Abstract: Objective: To describe the architecture and design rationale of meta-pipe, an open-source large language model (LLM)-agent pipeline that integrates th

model-releasesarxiv-cs-ai
30 Jun 2026
Safety

Metric Aggregation Divergence: A Hidden Validity Threat in Agent-Based Policy Optimization and a Contractual Remedy

DGX agent

arXiv:2606.29038v1 Announce Type: cross Abstract: Metric aggregation divergence (MAD) is the silent inconsistency that arises when distinct pipeline stages in an agent-based model coupled with a multi

safetyarxiv-cs-ai
30 Jun 2026
Model Releases

Multi-Agent Route Planning as a QUBO Problem

DGX agent

arXiv:2602.07913v2 Announce Type: replace Abstract: Multi-Agent Route Planning considers selecting vehicles, each associated with a single predefined route, such that route-level coverage utility is m

model-releasesarxiv-cs-ro
30 Jun 2026
Model Releases

OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks

DGX agent

arXiv:2606.29537v1 Announce Type: new Abstract: Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world computer use, limiting their ability to

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

When Does Overlap Help? OSU-Mem and a Cell-Conditional Analysis of Trajectory Memory for LLM Agents

DGX agent

arXiv:2606.28376v1 Announce Type: cross Abstract: Long-horizon large language model (LLM) agents accumulate interaction trajectories that quickly exceed any practical prompt budget, and existing memor

model-releasesarxiv-cs-ai
30 Jun 2026
Safety

ATOD: Annealed Turn-aware On-policy Distillation for Multi-turn Autonomous Agents

DGX agent

arXiv:2606.27814v1 Announce Type: new Abstract: Training small language-model agents for long-horizon interactive tasks requires both fast imitation and reward-driven improvement. On-policy distillati

safetyarxiv-cs-ai
29 Jun 2026
Model Releases

Multimodal Evaluator Preference Collapse: Cross-Modal Coupling in Self-Evolving Agents

DGX agent

arXiv:2606.16682v3 Announce Type: replace-cross Abstract: When AI agents use language models to evaluate their own outputs in a feedback loop, systematic biases emerge. We show that Evaluator Preferen

model-releasesarxiv-cs-cl
29 Jun 2026
Model Releases

QuantV2X: A Fully Quantized Multi-Agent System for Cooperative Perception

DGX agent

arXiv:2509.03704v2 Announce Type: replace Abstract: Cooperative perception through Vehicle-to-Everything (V2X) communication offers significant potential for enhancing vehicle perception by mitigating

model-releasesarxiv-cs-cv
29 Jun 2026
Safety

Diagnosing Task Insensitivity in Language Agents

DGX agent

arXiv:2606.26918v1 Announce Type: new Abstract: Large language models can serve as capable long-horizon agents, but their out-of-distribution (OOD) generalization remains weak. We identify a key sourc

safetyarxiv-cs-ai
26 Jun 2026
Hardware

EGG: An Expert-Guided Agent Framework for Kernel Generation

DGX agent

arXiv:2606.26758v1 Announce Type: new Abstract: High-performance GPU kernels are critical for reducing the exponentially growing computational costs of large language models (LLMs), but their developm

hardwarearxiv-cs-ai
26 Jun 2026
Safety

EVOM: Agentic Meta-Evolution of Actor-Critic Architectures for Reinforcement Learning

DGX agent

arXiv:2606.26327v1 Announce Type: cross Abstract: In actor-critic reinforcement learning, network architectures are typically manually designed. Automating this design is challenging because each cand

safetyarxiv-cs-ai
26 Jun 2026
Safety

Improving General Role-Playing Agents via Psychology-Grounded Reasoning and Role-Aware Policy Optimization

DGX agent

arXiv:2606.27025v1 Announce Type: new Abstract: Building general-purpose role-playing agents that faithfully portray any character from a natural-language profile remains challenging. The dominant par

safetyarxiv-cs-cl
26 Jun 2026
Model Releases

Memory Depth, Not Memory Access: Selective Parametric Consolidation for Long-Running Language Agents

DGX agent

arXiv:2606.26806v1 Announce Type: new Abstract: Long-running language agents need more than memory access. Retrieval systems can fetch past facts at query time, but they do not decide which experience

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

Post-Training Recipe, More Than Model Family, Shapes Multi-Agent LLM Conversational Behavior

DGX agent

arXiv:2606.20632v2 Announce Type: replace-cross Abstract: Multi-LLM systems use multiple language models to deliberate, judge each other's outputs, or coordinate as agents. Their value depends on the

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

R2D-RL: A RoboCup 2D Soccer Environment for Multi-Agent Reinforcement Learning

DGX agent

arXiv:2606.18786v2 Announce Type: replace Abstract: Robot soccer is a challenging testbed for multi-agent reinforcement learning because it combines partial observability, cooperative and adversarial

model-releasesarxiv-cs-ai
26 Jun 2026
Safety

Semantic Early-Stopping for Iterative LLM Agent Loops

DGX agent

arXiv:2606.27009v1 Announce Type: new Abstract: Multi-agent large language model (LLM) loops, for example a Writer that drafts and a Critic that revises, are almost always terminated by a fixed iterat

safetyarxiv-cs-ai
26 Jun 2026
Model Releases

When Actions Go Off-Task: Detecting and Correcting Misaligned Actions in Computer-Use Agents

DGX agent

arXiv:2602.08995v2 Announce Type: replace Abstract: Computer-use agents (CUAs) have made tremendous progress in the past year, yet they still frequently produce misaligned actions that deviate from th

model-releasesarxiv-cs-cl
26 Jun 2026
Safety

When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models

DGX agent

arXiv:2606.27288v1 Announce Type: new Abstract: Multi-model LLM systems such as routing, voting, cascades, fusion, and mixture-of-agents are used to beat single-model accuracy. We show that their gain

safetyarxiv-cs-ai
26 Jun 2026
Research

Where Do CoT Training Gains Land in LLM based Agents?

DGX agent

arXiv:2606.26935v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning is widely used in language-model agents, but prior work has shown that verbalized CoT is not always faithful and may in

researcharxiv-cs-ai
26 Jun 2026
Safety

Adaptive-Horizon Conflict-Based Search for Closed-Loop Multi-Agent Path Finding

DGX agent

arXiv:2602.12024v3 Announce Type: replace Abstract: Multi-Agent Path Finding (MAPF) is a core coordination problem for large robot fleets in automated warehouses and logistics. Existing approaches are

safetyarxiv-cs-ro
25 Jun 2026
Safety

ASAP: Agent-System Co-Design for Wall-Clock-Centered Auto HPO Research for ML Experiments

DGX agent

arXiv:2606.25207v1 Announce Type: cross Abstract: Hyperparameter Optimization (HPO) is essential for maximizing machine learning model performance, and its core challenge is sample efficiency: finding

safetyarxiv-cs-cl
25 Jun 2026
Safety

Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents

DGX agent

arXiv:2606.25852v1 Announce Type: new Abstract: Group-based reinforcement learning effectively post-trains LLM agents for long-horizon, sparse-reward tasks by deriving step-level credit from trajector

safetyarxiv-cs-lg
25 Jun 2026
Model Releases

Can Language Model Agents be Helpful Circuit Explainers in Mechanistic Interpretability?

DGX agent

arXiv:2606.24026v1 Announce Type: new Abstract: Mechanistic interpretability has made substantial progress in automatically localizing circuits, but explaining what localized components do remains lab

model-releasesarxiv-cs-ai
24 Jun 2026
Model Releases

DeepBD: A Grounded Agentic Workflow for Variant Prioritization and Diagnosis of Genetic Birth Defects

DGX agent

arXiv:2606.24779v1 Announce Type: cross Abstract: Birth defects are a major cause of fetal loss, neonatal morbidity and long-term disability. In the subset with suspected genetic etiologies, exome and

model-releasesarxiv-cs-ai
24 Jun 2026
Model Releases

Multimedia and Visual Analytics in the Agentic Era

DGX agent

arXiv:2504.06138v3 Announce Type: replace-cross Abstract: Professional users need tools to help them gain actionable insights from large multimedia collections. Foundation models and AI agents have ra

model-releasesarxiv-cs-ai
24 Jun 2026
Model Releases

Evo-RAD: Navigating Rare Retinal Disease Diagnosis via Self-Evolving Agentic Retrieval

DGX agent

arXiv:2606.22955v1 Announce Type: new Abstract: Large-scale pretrained foundation models have revolutionized general medical screening, but often falter on rare diseases because such conditions are un

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

Fara-1.5: Scalable Learning Environments for Computer Use Agents

DGX agent

arXiv:2606.20785v1 Announce Type: cross Abstract: Collecting computer use data from human demonstrations is expensive and slow, motivating the need for scalable generation strategies. This requires tw

model-releasesarxiv-cs-lg
23 Jun 2026
Safety

HEAS: Hierarchical Evolutionary Agent-Based Simulation Framework for Multi-Objective Policy Search

DGX agent

arXiv:2508.15555v4 Announce Type: replace-cross Abstract: HEAS is a Python framework that connects agent-based simulation, evolutionary search, and scenario-based evaluation in a single reproducible p

safetyarxiv-cs-lg
23 Jun 2026
Model Releases

Nous: A Predictive World Model for Long-Term Agent Memory

DGX agent

arXiv:2606.22030v1 Announce Type: cross Abstract: We present Nous, a novel agent memory architecture grounded in the principle that knowledge is prediction, not storage. Rather than persisting facts a

model-releasesarxiv-cs-lg
23 Jun 2026
Model Releases

Self-Evolution for Multi-Turn Tool-Calling Agents via Divergence-Point Preference Learning

DGX agent

arXiv:2606.23112v1 Announce Type: new Abstract: Multi-turn tool-using agents must coordinate long-horizon tool sequences while tracking dialogue state and policy constraints. Existing approaches often

model-releasesarxiv-cs-lg
23 Jun 2026
Safety

SkillHarness: Harnessing Safe Skills for Computer-Use Agents

DGX agent

arXiv:2606.20636v1 Announce Type: cross Abstract: Computer-Use Agents (CUAs) are increasingly deployed in dynamic interactive environments, creating a growing need for continual skill learning during

safetyarxiv-cs-lg
23 Jun 2026
Safety

HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation

DGX agent

arXiv:2606.11559v1 Announce Type: new Abstract: Reinforcement learning typically improves multi-turn agent capabilities through the terminal outcome of the trajectories, which makes it difficult to de

safetyarxiv-cs-ai
11 Jun 2026
Safety

MASK: Multi-Agent Semantic K-Scheduling for Risk-Sensitive 6G Robotics

DGX agent

arXiv:2606.11249v1 Announce Type: cross Abstract: Realizing the vision of 6G connected robotics requires reconciling high-performance collaborative control with the rigid spectral limitations of physi

safetyarxiv-cs-lg
11 Jun 2026
Local Ai

MODF-SIR: A Multi-agent Omni-modal Distilled Framework for Social Intelligence Reasoning

DGX agent

arXiv:2606.12018v1 Announce Type: new Abstract: We propose a multi-agent collaborative framework built upon a lightweight Multimodal Large Language Model (MLLM), specifically designed for social intel

local-aiarxiv-cs-ai
11 Jun 2026
Model Releases

A History-Aware Visually Grounded Critic for Computer Use Agents

DGX agent

arXiv:2606.11078v1 Announce Type: new Abstract: Various test-time interventions for Computer Use Agents (CUAs), including critic models, have been developed to improve performance through pre-executio

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

EEVEE: Towards Test-time Prompt Learning in the Real World for Self-Improving Agents

DGX agent

arXiv:2606.11182v1 Announce Type: cross Abstract: In this paper, we propose EEVEE, the first multi-dataset test-time prompt learning framework for LLM agents, enabling test-time prompt learning under

model-releasesarxiv-cs-ai
10 Jun 2026
← Previous
1…9899100101102…236
Next →