AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,289 results
Model Releases

Effective Reinforcement Learning for Agentic Search by Recycling Zero-Variance Queries During Training

DGX agent

arXiv:2606.10709v1 Announce Type: cross Abstract: The use of GRPO-style algorithms has become the standard strategy for training LLM search agents under outcome-only rewards. With these algorithms, a

model-releasesarxiv-cs-ai
10 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Human-AI Coordination Zones: A Framework for Designing Human-in-the-Loop Experiences with Agentic AI

DGX agent

arXiv:2606.09848v1 Announce Type: cross Abstract: As generative and agentic AI becomes embedded in everyday products, practitioners face a persistent challenge: how to design human-AI coordination --

safetyarxiv-cs-ai
10 Jun 2026
Model Releases

Less Context, More Accuracy: A Bi-Temporal Memory Engine for LLM Agents Where a Lean Retrieved Context Beats the Full History

DGX agent

arXiv:2606.09900v1 Announce Type: cross Abstract: Long-term memory is the missing layer for LLM agents: across sessions they forget, and the common workaround -- replaying the whole history into the p

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Moonshine: An Autonomous Mathematical Research Agent Centered on Conjecture Generation

DGX agent

arXiv:2606.10806v1 Announce Type: new Abstract: Moonshine is an autonomous agent whose central objective is to generate mathematical conjectures. Its core capability is to extract structure from class

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

What Fits (Into Few Tokens) Doesn't Overfit: Compression and Generalization in ML Research Agents

DGX agent

arXiv:2606.11045v1 Announce Type: new Abstract: Reusing a held-out benchmark adaptively should, in principle, invite overfitting. Yet benchmark-driven machine learning (ML) has produced surprisingly l

model-releasesarxiv-cs-ai
10 Jun 2026
Safety

Can the Environment Speak for Itself? T^{2}-GRPO: A Turn-Trajectory Group Relative Policy Optimization for Caregiver Agents

DGX agent

arXiv:2606.08875v1 Announce Type: new Abstract: Optimizing large language models (LLMs) for long-horizon caregiver agents requires balancing delayed task objectives with immediate environment dynamics

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

Rosetta Memory: Adaptive Memory for Cross-LLM Agents

DGX agent

arXiv:2606.07711v1 Announce Type: cross Abstract: Memory is the key component for transforming a stateless LLM into a persistent, evolving agent through experience accumulation, long-horizon planning,

model-releasesarxiv-cs-ai
9 Jun 2026
Safety

Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety

DGX agent

arXiv:2606.06529v1 Announce Type: new Abstract: An attacker that strategically chooses when to attack is much harder to catch than one that attacks indiscriminately. AI control is a safety framework f

safetyarxiv-cs-ai
8 Jun 2026
Model Releases

Autonomous heterogeneous catalyst discovery with a self-evolving multi-agent digital twin

DGX agent

arXiv:2606.05050v1 Announce Type: cross Abstract: Theoretical heterogeneous catalysis promises rapid catalyst discovery, yet computational and machine-learning predictions often deviate from experimen

model-releasesarxiv-cs-ai
8 Jun 2026
Safety

MADRAG: Multi-Agent Debate with Retrieval-Augmented Generation for Training-Free Analytic Essay Scoring

DGX agent

arXiv:2606.06754v1 Announce Type: cross Abstract: We present MADRAG, a training-free framework for analytic essay scoring that combines multi-agent reasoning with retrieval-augmented grounding. Unlike

safetyarxiv-cs-cl
8 Jun 2026
Model Releases

MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism

DGX agent

arXiv:2606.07512v1 Announce Type: cross Abstract: Current Vision-Language Models struggle with hours-long videos because processing full-length visual sequences induces prohibitive token explosion and

model-releasesarxiv-cs-ai
8 Jun 2026
Safety

SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating

DGX agent

arXiv:2606.07074v1 Announce Type: cross Abstract: Deep research agents have demonstrated remarkable capabilities in complex information-seeking tasks, yet this power comes at a steep computational cos

safetyarxiv-cs-ai
8 Jun 2026
Model Releases

CangLing-KnowFlow: A Unified Knowledge-and-Flow-fused Agent for Comprehensive Remote Sensing Applications

DGX agent

arXiv:2512.15231v3 Announce Type: replace Abstract: The automated and intelligent processing of massive remote sensing (RS) datasets is critical in Earth observation (EO). Existing automated systems a

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Critic-Guided Heterogeneous Multi-Agent Reasoning for Reliable Mathematical Problem Solving

DGX agent

arXiv:2606.05704v1 Announce Type: new Abstract: Recent Large Language Models (LLMs) have shown impressive reasoning abilities; but they are still susceptible to hallucinations, intermediate reasoning

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Data Flow Control: Data Safety Policies for AI Agents

DGX agent

arXiv:2606.05679v1 Announce Type: cross Abstract: Agents increasingly generate SQL, orchestrate pipelines, and automate data analysis on behalf of users. While recent work improves query correctness,

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

TAPO: Tool-Aware Policy Optimization via Credit Transfer for Multimodal Search Agents

DGX agent

arXiv:2606.05784v1 Announce Type: new Abstract: We identify and formally characterize credit misassignment as a systematic failure mode of GRPO in tool-augmented multimodal search agents: its uniform

model-releasesarxiv-cs-ai
6 Jun 2026
Local Ai

TOKI: A Bitemporal Operator Algebra for Contradiction Resolution in LLM-Agent Persistent Memory

DGX agent

arXiv:2606.06240v1 Announce Type: cross Abstract: Persistent memory for an LLM agent is a write-heavy substrate: every belief update is a versioned write, and a new claim may contradict a stored one.

local-aiarxiv-cs-ai
6 Jun 2026
Model Releases

ToolChoiceConfusion: Causal Minimal Tool Filtering for Reliable LLM Agents

DGX agent

arXiv:2606.06284v1 Announce Type: new Abstract: Large language model agents increasingly rely on external tools, but larger tool menus can reduce reliability and efficiency by increasing wrong-tool ca

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

When Should Memory Stay Silent: Measuring Memory-Use Boundaries in Memory-Augmented Conversational Agents

DGX agent

arXiv:2606.06055v1 Announce Type: new Abstract: Long-term memory enables language model agents to support personalized interactions, but it remains unclear when available memories warrant integration

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

When Tools Fail: Benchmarking Dynamic Replanning and Anomaly Recovery in LLM Agents

DGX agent

arXiv:2606.05806v1 Announce Type: new Abstract: Existing benchmarks evaluate Tool-Integrated Reasoning (TIR) in LLMs on idealized ''happy paths'', largely overlooking real-world tool failures. We intr

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time?

DGX agent

arXiv:2606.05553v1 Announce Type: new Abstract: Role-playing language agents (RPLAs) should play characters whose values and behavior evolve as the story progresses, not maintain a fixed persona. Exis

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Asuka-Bench: Benchmarking Code Agents on Underspecified User Intent and Multi-Round Refinement

DGX agent

arXiv:2606.05920v1 Announce Type: cross Abstract: Existing code-generation benchmarks score a single mapping from a complete prompt to a one-shot output. However, real web development is different. Us

model-releasesarxiv-cs-cl
5 Jun 2026
Safety

EMBER: Efficient Memory via Budgeted Evidence Retention for Long-Horizon Agents

DGX agent

arXiv:2606.05894v1 Announce Type: new Abstract: Long-horizon agents can archive large histories, but future answers still incur retrieval, rereading, and context costs. When retained memory misses ans

safetyarxiv-cs-cl
5 Jun 2026
Safety

Toward Culturally Aligned LLMs through Ontology-Guided Multi-Agent Reasoning

DGX agent

arXiv:2601.21700v3 Announce Type: replace Abstract: Large Language Models (LLMs) increasingly support culturally sensitive decision making, yet often exhibit misalignment due to skewed pretraining dat

safetyarxiv-cs-cl
5 Jun 2026
Safety

When Evidence is Sparse: Weakly Supervised Early Failure Alerting in Dialogs and LLM-Agent Trajectories

DGX agent

arXiv:2606.05414v1 Announce Type: new Abstract: Early failure alerting requires deciding, while a dialog or agent trajectory is still unfolding, whether to flag it as likely to fail. This is challengi

safetyarxiv-cs-cl
5 Jun 2026
Model Releases

Adaptive Minds: Empowering Agents with LoRA-as-Tools

DGX agent

arXiv:2510.15416v2 Announce Type: replace Abstract: We investigate a framework in which LoRA adapters are treated as callable tools that a base language model can dynamically select and invoke. We hyp

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

Caught in the Act(ivation): Toward Pre-Output and Multi-Turn Detection of Credential Exfiltration by LLM Agents

DGX agent

arXiv:2606.04141v1 Announce Type: cross Abstract: LLM agents often place sensitive credentials in the same context window as untrusted retrieved content, creating a direct path for indirect prompt inj

model-releasesarxiv-cs-ai
4 Jun 2026
Local Ai

ARBOR: Online Process Rewards via a Reusable Rubric Buffer for Search Agents

DGX agent

arXiv:2606.03239v1 Announce Type: new Abstract: LLM-based search agents are trained predominantly with outcome-only reward, leaving the search process itself unsupervised. This signal degenerates on o

local-aiarxiv-cs-cl
3 Jun 2026
Model Releases

Assistax: A Multi-Agent Hardware-Accelerated Reinforcement Learning Benchmark for Assistive Robotics

DGX agent

arXiv:2507.21638v2 Announce Type: replace Abstract: The development of reinforcement learning (RL) algorithms has been largely driven by ambitious challenge tasks and benchmarks. Games have dominated

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

EvoDrive: Pareto Evolution for Safety-Critical Autonomous Driving via Self-Improving LLM Agents

DGX agent

arXiv:2606.03678v1 Announce Type: new Abstract: Generating safety-critical scenarios is essential for validating and improving autonomous driving systems, yet it inherently requires maximizing adversa

model-releasesarxiv-cs-ai
3 Jun 2026
Safety

Margin Play: A Multi-Agent System For Public Policy Analysis In The Brazilian Equatorial Margin

DGX agent

arXiv:2606.02614v1 Announce Type: cross Abstract: The Brazilian Equatorial Margin (BEM) is Brazil's next offshore oil frontier, with operations expected to begin in 2026 in the Foz do Amazonas basin.

safetyarxiv-cs-ai
3 Jun 2026
Hardware

MOSAIC: Efficient Mixture-of-Agent Scheduling via Adaptive Aggregation and Inference Concurrency

DGX agent

arXiv:2606.03014v1 Announce Type: new Abstract: Mixture-of-Agents (MoA) systems improve reasoning accuracy by routing each query to multiple expert LLMs and aggregating their outputs. Efficiently exec

hardwarearxiv-cs-lg
3 Jun 2026
Research

RGMem: Renormalization Group-inspired Memory Evolution for Language Agents

DGX agent

arXiv:2510.16392v3 Announce Type: replace Abstract: Personalized and continuous interactions are critical for LLM-based conversational agents, yet finite context windows and static parametric memory h

researcharxiv-cs-ai
3 Jun 2026
Safety

Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning

DGX agent

arXiv:2606.03762v1 Announce Type: cross Abstract: Agentic reinforcement learning (RL) equips large language models (LLMs) with tool-use capabilities that substantially improve reasoning on complex tas

safetyarxiv-cs-ai
3 Jun 2026
Model Releases

TSQAgent: Rating Time Series Data Quality via Dedicated Agentic Reasoning

DGX agent

arXiv:2606.03629v1 Announce Type: new Abstract: Assessing the quality of time series (TS) data is fundamental yet inherently challenging due to the multifaceted nature of quality dimensions. Recently,

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Cross-Environment Neural Reranking for Sample-Efficient Action Selection in Text-Based Agents

DGX agent

arXiv:2606.02204v1 Announce Type: new Abstract: Large language model agents achieve strong performance on text-based benchmarks but incur prohibitive inference costs, motivating the use of compact neu

model-releasesarxiv-cs-cl
2 Jun 2026
Research

Doing What They Say, Not What They Reason: Locating the Faithfulness Gap in LLM Agents

DGX agent

arXiv:2606.00476v1 Announce Type: new Abstract: Do LLM agents act on the reasoning they state? This question of process fidelity is central to using LLMs in social simulation, yet it is hard to measur

researcharxiv-cs-ai
2 Jun 2026
Safety

Leyline: KV Cache Directives for Agentic Inference

DGX agent

arXiv:2606.01065v1 Announce Type: cross Abstract: Modern KV cache management assumes the chatbot workload: prompts arrive once and the cache grows append-only, so prefix caching and forward-only evict

safetyarxiv-cs-ai
2 Jun 2026
Model Releases

LLM Consortium for Software Design Refinement: A Controlled Experiment on Multi-Agent Collaboration Topologies

DGX agent

arXiv:2606.01490v1 Announce Type: cross Abstract: We present a controlled experiment evaluating 12 multi-agent LLM collaboration topologies for software architecture design. Using a 2imes2imes2 factor

model-releasesarxiv-cs-ai
2 Jun 2026
Safety

MobEvolve: An Agentic Self-Evolving Heuristic System for Interpretable Human Mobility Generation

DGX agent

arXiv:2606.01640v1 Announce Type: new Abstract: Human mobility generation aims to synthesize realistic trip chains for target populations based on individual features. Existing paradigms, including de

safetyarxiv-cs-ai
2 Jun 2026
Hardware

Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving

DGX agent

arXiv:2606.01839v1 Announce Type: cross Abstract: LLM-based agents resolve a user task through many turns of dependent inference and tool calls, producing a workload whose total cost is unknown when t

hardwarearxiv-cs-lg
2 Jun 2026
Safety

On Effectiveness and Efficiency of Agentic Tool-calling and RL Training

DGX agent

arXiv:2606.00135v1 Announce Type: cross Abstract: Tool-calling is a central component of modern large language model (LLM) agents, equipping them with skills beyond their parametric knowledge. This pa

safetyarxiv-cs-ai
2 Jun 2026
Model Releases

RescueBench: Can Embodied Agents Save Lives in the Wild ?

DGX agent

arXiv:2606.01848v1 Announce Type: new Abstract: Search-and-rescue (SAR) requires embodied agents to explore unfamiliar environments under multimodal uncertainty, perform multi-stage interactions, and

model-releasesarxiv-cs-cv
2 Jun 2026
Safety

Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents

DGX agent

arXiv:2605.30590v1 Announce Type: cross Abstract: Two clinical AI systems can score nearly identically on coverage-based rubrics yet behave radically differently when their patient inputs change: one

safetyarxiv-cs-ai
1 Jun 2026
Model Releases

DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories

DGX agent

arXiv:2602.10809v2 Announce Type: replace Abstract: Existing multimodal retrieval systems excel at semantic matching but implicitly assume that query-image relevance can be measured in isolation. This

model-releasesarxiv-cs-cv
1 Jun 2026
Safety

Detect in Any Scene: An Agentic Framework for Object Detection with Experience-Aware Reasoning

DGX agent

arXiv:2605.31174v1 Announce Type: new Abstract: Object detection in real-world scenarios remains challenging due to diverse image degradations and heterogeneous object distributions, which significant

safetyarxiv-cs-cv
1 Jun 2026
Model Releases

Eywa: Provenance-Grounded Long-Term Memory for AI Agents

DGX agent

arXiv:2605.30771v1 Announce Type: new Abstract: AI agents that persist across sessions need memory they can retrieve, audit, update, and erase. Existing memory systems often collapse source evidence,

model-releasesarxiv-cs-cl
1 Jun 2026
Research

Post-Training LLMs as Better Decision-Making Agents: A Regret-Minimization Approach

DGX agent

arXiv:2511.04393v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as 'agents' for decision-making (DM) in interactive and dynamic environments. Yet, since they

researcharxiv-cs-ai
1 Jun 2026
← Previous
1…99100101102103…236
Next →