AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,289 results
Model Releases

GVGAI-LLM: Evaluating Large Language Model Agents with Infinite Games

DGX agent

arXiv:2508.08501v3 Announce Type: replace Abstract: We introduce GVGAI-LLM, a video game benchmark for evaluating the reasoning and problem-solving capabilities of large language models (LLMs). Built

model-releasesarxiv-cs-ai
19 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

MirrorBench: A Benchmark to Evaluate Conversational User-Proxy Agents for Human-Likeness

DGX agent

arXiv:2601.08118v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used as human simulators, both for evaluating conversational systems and for generating fine-tuning da

model-releasesarxiv-cs-ai
19 May 2026
Safety

NEWTON: Agentic Planning for Physically Grounded Video Generation

DGX agent

arXiv:2605.18396v1 Announce Type: new Abstract: Video generation models produce visually compelling results but systematically violate physical commonsense -- on VideoPhy-2, the best model achieves on

safetyarxiv-cs-cv
19 May 2026
Safety

Whispers in the Noise: Surrogate-Guided Concept Awakening via a Multi-Agent Framework

DGX agent

arXiv:2605.18150v1 Announce Type: new Abstract: Diffusion models (DMs) are widely used for text-to-image generation, but their strong generative capabilities also raise concerns about unsafe or undesi

safetyarxiv-cs-ai
19 May 2026
Model Releases

Confirming Correct, Missing the Rest: LLM Tutoring Agents Struggle Where Feedback Matters Most

DGX agent

arXiv:2605.16207v1 Announce Type: new Abstract: Effective tutoring requires distinguishing optimal, valid but suboptimal, and incorrect student solutions, a distinction central to intelligent tutoring

model-releasesarxiv-cs-ai
18 May 2026
Research

Effective Harness Engineering for Algorithm Discovery with Coding Agents

DGX agent

arXiv:2605.15221v1 Announce Type: cross Abstract: AlphaEvolve and FunSearch have demonstrated the potential of combining large language models (LLMs) with evolutionary search for automated algorithm d

researcharxiv-cs-ai
18 May 2026
Safety

FinReporting: An Agentic Workflow for Localized Reporting of Cross-Jurisdiction Financial Disclosures

DGX agent

arXiv:2604.05966v2 Announce Type: replace Abstract: Financial reporting systems increasingly leverage Large Language Models (LLMs) to extract and summarize corporate disclosures. However, most existin

safetyarxiv-cs-cl
18 May 2026
Model Releases

Hidden in Memory: Sleeper Memory Poisoning in LLM Agents

DGX agent

arXiv:2605.15338v1 Announce Type: cross Abstract: Large language models are increasingly augmented with persistent memory, allowing assistants to store user-specific information across sessions for pe

model-releasesarxiv-cs-ai
18 May 2026
Agents

On the Convergence Rates of Federated Q-Learning across Heterogeneous Environments

DGX agent

arXiv:2409.03897v3 Announce Type: replace Abstract: Large-scale multi-agent systems are often deployed across wide geographic areas, where agents interact with heterogeneous environments. There is an

agentsarxiv-cs-lg
18 May 2026
Model Releases

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation

DGX agent

arXiv:2605.16079v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have shown significant progress in video understanding, yet they face substantial challenges in tasks requiring p

model-releasesarxiv-cs-ai
18 May 2026
Model Releases

Agentic Systems as Boosting Weak Reasoning Models

DGX agent

arXiv:2605.14163v1 Announce Type: new Abstract: Can a committee of weak reasoning-model calls reach the performance of much stronger models? We study verifier-backed committee search as inference-time

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Distribution-Aware Algorithm Design with LLM Agents

DGX agent

arXiv:2605.14141v1 Announce Type: new Abstract: We study learning when the learned object is executable solver code rather than a predictor. In this setting, correctness is not enough: two solvers may

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Mini-JEPA Foundation Model Fleet Enables Agentic Hydrologic Intelligence

DGX agent

arXiv:2605.14120v1 Announce Type: cross Abstract: Geospatial foundation models compress multispectral observations into dense embeddings increasingly used in natural-language environmental reasoning s

model-releasesarxiv-cs-cl
15 May 2026
Research

Watermarking Game-Playing Agents in Perfect-Information Extensive-Form Games

DGX agent

arXiv:2605.14283v1 Announce Type: cross Abstract: Watermarking techniques for large language models (LLMs), which encode hidden information in the output so its source can be verified, have gained sig

researcharxiv-cs-ai
15 May 2026
Model Releases

RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation

DGX agent

arXiv:2605.13542v1 Announce Type: new Abstract: Intensive care units (ICU) generate long, dense and evolving streams of clinical information, where physicians must repeatedly reassess patient states u

model-releasesarxiv-cs-ai
14 May 2026
Agents

TRIAGE: Evaluating Prospective Metacognitive Control in LLMs under Resource Constraints

DGX agent

arXiv:2605.13414v1 Announce Type: new Abstract: Deploying language models as autonomous agents requires more than per-task accuracy: when an agent faces a queue of problems under a finite token budget

agentsarxiv-cs-ai
14 May 2026
Model Releases

From Web to Pixels: Bringing Agentic Search into Visual Perception

DGX agent

arXiv:2605.12497v1 Announce Type: new Abstract: Visual perception connects high-level semantic understanding to pixel-level perception, but most existing settings assume that the decisive evidence for

model-releasesarxiv-cs-cv
13 May 2026
Agents

Joint Learning of Hierarchical Neural Options and Abstract World Model

DGX agent

arXiv:2602.02799v2 Announce Type: replace Abstract: Building agents that can perform new skills by composing existing skills is a long-standing goal of AI agent research. Towards this end, we investig

agentsarxiv-cs-lg
13 May 2026
Model Releases

PresentAgent-2: Towards Generalist Multimodal Presentation Agents

DGX agent

arXiv:2605.11363v1 Announce Type: cross Abstract: Presentation generation is moving beyond static slide creation toward end-to-end presentation video generation with research grounding, multimodal med

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Test-Time Compute for Dense Retrieval: Agentic Program Generation with Frozen Embedding Models

DGX agent

arXiv:2605.11374v1 Announce Type: cross Abstract: Test-time compute is widely believed to benefit only large reasoning models. We show it also helps small embedding models. Most modern embedding check

model-releasesarxiv-cs-cl
13 May 2026
Safety

AI-Care: A Conversational Agentic System for Task Coordination in Alzheimer's Disease Care

DGX agent

arXiv:2605.08480v1 Announce Type: new Abstract: Individuals with Alzheimer's disease (AD) and Alzheimer's disease-related dementia (ADRD) experience memory and thinking changes that impact their abili

safetyarxiv-cs-ai
12 May 2026
Model Releases

AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents

DGX agent

arXiv:2605.10876v1 Announce Type: cross Abstract: Recent advances in machine learning and large-scale biological data collections have revived the prospect of building a virtual cell, a computational

model-releasesarxiv-cs-ai
12 May 2026
Safety

Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning

DGX agent

arXiv:2605.09395v1 Announce Type: new Abstract: In this paper, we propose the first VLnderline{extbf{M}} nderline{extbf{a}}gentic nderline{extbf{r}}easoning framework for few-nderline{extbf{s}}hot mul

safetyarxiv-cs-ai
12 May 2026
Safety

Learning the Preferences of a Learning Agent

DGX agent

arXiv:2605.09217v1 Announce Type: new Abstract: For AI systems to be useful to humans, they must understand and act in accordance with our values and preferences. Since specifying preferences is a har

safetyarxiv-cs-ai
12 May 2026
Model Releases

Source or It Didn't Happen: A Multi-Agent Framework for Citation Hallucination Detection

DGX agent

arXiv:2605.08583v1 Announce Type: new Abstract: Large language models are increasingly used in scientific writing, yet they can fabricate citation-shaped references that appear plausible but fail bibl

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

A^2RD: Agentic Autoregressive Diffusion for Long Video Consistency

DGX agent

arXiv:2605.06924v1 Announce Type: cross Abstract: Synthesizing consistent and coherent long video remains a fundamental challenge. Existing methods suffer from semantic drift and narrative collapse ov

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

MAVEN: Multi-Agent Verification-Elaboration Network with In-Step Epistemic Auditing

DGX agent

arXiv:2605.07646v1 Announce Type: cross Abstract: While explicit reasoning trajectories enhance model interpretability, existing paradigms often rely on monolithic chains that lack intermediate verifi

model-releasesarxiv-cs-ai
11 May 2026
Safety

PACEvolve++: Improving Test-time Learning for Evolutionary Search Agents

DGX agent

arXiv:2605.07039v1 Announce Type: new Abstract: Large language models have become drivers of evolutionary search, but most systems rely on a fixed, prompt-elicited policy to sample next candidates. Th

safetyarxiv-cs-lg
11 May 2026
Safety

Agent-Based Modeling of Low-Emission Fertilizer Adoption for Dairy Farm Decarbonisation using Empirical Farm Data

DGX agent

arXiv:2605.03648v1 Announce Type: new Abstract: To understand complex system dynamics in dairy farming, it is essential to use modeling tools that capture farm heterogeneity, social interactions, and

safetyarxiv-cs-ai
7 May 2026
Model Releases

Physics-Grounded Multi-Agent Architecture for Traceable, Risk-Aware Human-AI Decision Support in Manufacturing

DGX agent

arXiv:2605.04003v1 Announce Type: cross Abstract: High-precision CNC machining of free-form aerospace components requires bounded compensations informed by inspection, simulation, and process knowledg

model-releasesarxiv-cs-ai
7 May 2026
Model Releases

VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA

DGX agent

arXiv:2605.04870v1 Announce Type: new Abstract: Video text-based visual question answering (Video TextVQA) aims to answer questions by reasoning over visual textual content appearing in videos. Despit

model-releasesarxiv-cs-cv
7 May 2026
Model Releases

CreativityBench: Evaluating Agent Creative Reasoning via Affordance-Based Tool Repurposing

DGX agent

arXiv:2605.02910v1 Announce Type: cross Abstract: Recent advances in large language models have led to strong performance on reasoning and environment-interaction tasks, yet their ability for creative

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

CuraView: A Multi-Agent Framework for Medical Hallucination Detection with GraphRAG-Enhanced Knowledge Verification

DGX agent

arXiv:2605.03476v1 Announce Type: new Abstract: Discharge summaries require extracting critical information from lengthy electronic health records (EHRs), a process that is labor-intensive when perfor

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

Toward a Science of Intent: Closure Gaps and Delegation Envelopes for Open-World AI Agents

DGX agent

arXiv:2604.25000v2 Announce Type: replace Abstract: Recent work has framed intelligence in verifiable tasks as reducing time-to-solution through learned structure and test-time search, while systems w

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Towards Agentic Runtime Healing

DGX agent

arXiv:2408.01055v2 Announce Type: replace-cross Abstract: Self-healing systems have long been a focus of research, aiming to enable software to recover from unexpected runtime errors without human int

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

ARA: Agentic Reproducibility Assessment For Scalable Support Of Scientific Peer-Review

DGX agent

arXiv:2605.02651v1 Announce Type: cross Abstract: Scientific peer review increasingly struggles to assess reproducibility at the scale and complexity of modern research output. Evaluating reproducibil

model-releasesarxiv-cs-lg
5 May 2026
Model Releases

GAZE: Grounded Agentic Zero-shot Evaluation with Viewer-Level Tools and Literature Retrieval on Rare Brain MRI

DGX agent

arXiv:2605.00876v1 Announce Type: cross Abstract: Vision-language models (VLMs) read an image and produce text in a single forward pass, whereas radiologists typically inspect an image several times a

model-releasesarxiv-cs-cv
5 May 2026
Applications

Graph Query Generation with Constraint-guided Large Language Agents

DGX agent

arXiv:2605.00845v1 Announce Type: cross Abstract: Knowledge Graph Question Answering (KGQA) has advanced through structured query generation, yet most efforts target RDF/SPARQL, leaving Cypher and pro

applicationsarxiv-cs-cl
5 May 2026
Research

MemORAI: Memory Organization and Retrieval via Adaptive Graph Intelligence for LLM Conversational Agents

DGX agent

arXiv:2605.01386v1 Announce Type: new Abstract: Large Language Models (LLMs) lack persistent memory for long-term personalized conversations. Existing graph-based memory systems suffer from informatio

researcharxiv-cs-cl
5 May 2026
Safety

Separation Assurance between Heterogeneous Fleets of Small Unmanned Aerial Systems via Multi-Agent Reinforcement Learning

DGX agent

arXiv:2605.01041v1 Announce Type: cross Abstract: In the envisioned future dense urban airspace, multiple companies will operate heterogeneous fleets of small unmanned aerial systems (sUASs), where ea

safetyarxiv-cs-lg
5 May 2026
Applications

A Pattern Language for Resilient Visual Agents

DGX agent

arXiv:2604.28001v1 Announce Type: new Abstract: Integrating multimodal foundation models into enterprise ecosystems presents a fundamental software architecture challenge. Architects must balance comp

applicationsarxiv-cs-ai
1 May 2026
Model Releases

CareGuardAI: Context-Aware Multi-Agent Guardrails for Clinical Safety & Hallucination Mitigation in Patient-Facing LLMs

DGX agent

arXiv:2604.26959v1 Announce Type: cross Abstract: Integrating large language models (LLMs) into patient-facing healthcare systems offers significant potential to improve access to medical information.

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

Grounding Agent Memory in Contextual Intent

DGX agent

arXiv:2601.10702v2 Announce Type: replace-cross Abstract: Deploying large language models in long-horizon, goal-oriented interactions remains challenging because similar entities and facts recur under

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

RoadMapper: A Multi-Agent System for Roadmap Generation of Solving Complex Research Problems

DGX agent

arXiv:2604.27616v1 Announce Type: new Abstract: People commonly leverage structured content to accelerate knowledge acquisition and research problem solving. Among these, roadmaps guide researchers th

model-releasesarxiv-cs-cl
1 May 2026
Model Releases

A self-evolving agent for explainable diagnosis of DFT-experiment band-gap mismatch

DGX agent

arXiv:2604.26703v1 Announce Type: cross Abstract: Standard density functional theory (DFT) routinely misclassifies the electronic ground state of correlated and structurally complex compounds, predict

model-releasesarxiv-cs-ai
30 Apr 2026
Agents

Diagnosis, Bad Planning & Reasoning. Treatment, SCOPE -- Planning for Hybrid Querying over Clinical Trial Data

DGX agent

arXiv:2604.25120v1 Announce Type: new Abstract: We study clinical trial table reasoning, where answers are not directly stored in visible cells but must be reasoned from semantic understanding through

agentsarxiv-cs-cl
29 Apr 2026
Model Releases

DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios

DGX agent

arXiv:2604.25914v1 Announce Type: new Abstract: Real-world data visualization (DV) requires native environmental grounding, cross-platform evolution, and proactive intent alignment. Yet, existing benc

model-releasesarxiv-cs-cl
29 Apr 2026
Agents

MiMo-Embodied: X-Embodied Foundation Model Technical Report

DGX agent

arXiv:2511.16518v2 Announce Type: replace-cross Abstract: We open-source MiMo-Embodied, the first cross-embodied foundation model to successfully integrate and achieve state-of-the-art performance in

agentsarxiv-cs-cl
29 Apr 2026
← Previous
1…114115116117118…236
Next →