AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,289 results
Model Releases

Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization

DGX agent

arXiv:2604.12290v1 Announce Type: new Abstract: Current LLM agent benchmarks, which predominantly focus on binary pass/fail tasks such as code generation or search-based question answering, often negl

model-releasesarxiv-cs-ai
15 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning

DGX agent

arXiv:2601.06794v2 Announce Type: replace Abstract: Critique-guided reinforcement learning (RL) has emerged as a powerful paradigm for training LLM agents by augmenting sparse outcome rewards with nat

safetyarxiv-cs-ai
15 Apr 2026
Model Releases

Policy-Invisible Violations in LLM-Based Agents

DGX agent

arXiv:2604.12177v1 Announce Type: new Abstract: LLM-based agents can execute actions that are syntactically valid, user-sanctioned, and semantically appropriate, yet still violate organizational polic

model-releasesarxiv-cs-ai
15 Apr 2026
Model Releases

Spatial Atlas: Compute-Grounded Reasoning for Spatial-Aware Research Agent Benchmarks

DGX agent

arXiv:2604.12102v1 Announce Type: new Abstract: We introduce compute-grounded reasoning (CGR), a design paradigm for spatial-aware research agents in which every answerable sub-problem is resolved by

model-releasesarxiv-cs-ai
15 Apr 2026
Model Releases

Thought-Retriever: Don't Just Retrieve Raw Data, Retrieve Thoughts for Memory-Augmented Agentic Systems

DGX agent

arXiv:2604.12231v1 Announce Type: new Abstract: Large language models (LLMs) have transformed AI research thanks to their powerful internal capabilities and knowledge. However, existing LLMs still fai

model-releasesarxiv-cs-cl
15 Apr 2026
Model Releases

Towards Long-horizon Agentic Multimodal Search

DGX agent

arXiv:2604.12890v1 Announce Type: cross Abstract: Multimodal deep search agents have shown great potential in solving complex tasks by iteratively collecting textual and visual evidence. However, mana

model-releasesarxiv-cs-ai
15 Apr 2026
Safety

A Dual-Positive Monotone Parameterization for Multi-Segment Bids and a Validity Assessment Framework for Reinforcement Learning Agent-based Simulation of Electricity Markets

DGX agent

arXiv:2604.10252v1 Announce Type: new Abstract: Reinforcement learning agent-based simulation (RL-ABS) has become an important tool for electricity market mechanism analysis and evaluation. In the mod

safetyarxiv-cs-ai
14 Apr 2026
Model Releases

From Translation to Superset: Benchmark-Driven Evolution of a Production AI Agent from Rust to Python

DGX agent

arXiv:2604.11518v1 Announce Type: cross Abstract: Cross-language migration of large software systems is a persistent engineering challenge, particularly when the source codebase evolves rapidly. We pr

model-releasesarxiv-cs-ai
14 Apr 2026
Safety

Hubble: An LLM-Driven Agentic Framework for Safe and Automated Alpha Factor Discovery

DGX agent

arXiv:2604.09601v1 Announce Type: new Abstract: Discovering predictive alpha factors in quantitative finance remains a formidable challenge due to the vast combinatorial search space and inherently lo

safetyarxiv-cs-ai
14 Apr 2026
Agents

Improving Layout Representation Learning Across Inconsistently Annotated Datasets via Agentic Harmonization

DGX agent

arXiv:2604.11042v1 Announce Type: new Abstract: Fine-tuning object detection (OD) models on combined datasets assumes annotation compatibility, yet datasets often encode conflicting spatial definition

agentsarxiv-cs-cv
14 Apr 2026
Agents

Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards

DGX agent

arXiv:2604.09855v1 Announce Type: new Abstract: The recent advancement of Large Language Models (LLMs) has established their potential as autonomous interactive agents. However, they often struggle in

agentsarxiv-cs-ai
14 Apr 2026
Safety

MADQRL: Distributed Quantum Reinforcement Learning Framework for Multi-Agent Environments

DGX agent

arXiv:2604.11131v1 Announce Type: new Abstract: Reinforcement learning (RL) is one of the most practical ways to learn from real-life use-cases. Motivated from the cognitive methods used by humans mak

safetyarxiv-cs-ai
14 Apr 2026
Model Releases

Multi-ORFT: Stable Online Reinforcement Fine-Tuning for Multi-Agent Diffusion Planning in Cooperative Driving

DGX agent

arXiv:2604.11734v1 Announce Type: cross Abstract: Closed-loop cooperative driving requires planners that generate realistic multimodal multi-agent trajectories while improving safety and traffic effic

model-releasesarxiv-cs-ai
14 Apr 2026
Agents

OpeFlo: Automated UX Evaluation via Simulated Human Web Interaction with GUI Grounding

DGX agent

arXiv:2604.09581v1 Announce Type: new Abstract: Evaluating web usability typically requires time-consuming user studies and expert reviews, which often limits iteration speed during product developmen

agentsarxiv-cs-ai
14 Apr 2026
Model Releases

Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind

DGX agent

arXiv:2604.11666v1 Announce Type: cross Abstract: As large language models (LLMs) become the engine behind conversational systems, their ability to reason about the intentions and states of their dial

model-releasesarxiv-cs-ai
14 Apr 2026
Agents

The Devil is in the Details -- From OCR for Old Church Slavonic to Purely Visual Stemma Reconstruction

DGX agent

arXiv:2604.11724v1 Announce Type: new Abstract: The age of artificial intelligence has brought many new possibilities and pitfalls in many fields and tasks. The devil is in the details, and those come

agentsarxiv-cs-cv
14 Apr 2026
Model Releases

Time is Not a Label: Continuous Phase Rotation for Temporal Knowledge Graphs and Agentic Memory

DGX agent

arXiv:2604.11544v1 Announce Type: cross Abstract: Structured memory representations such as knowledge graphs are central to autonomous agents and other long-lived systems. However, most existing appro

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

AgentCE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments

DGX agent

arXiv:2604.06111v2 Announce Type: replace Abstract: Existing Agent benchmarks suffer from two critical limitations: high environment interaction overhead (up to 41% of total evaluation time) and imbal

model-releasesarxiv-cs-ai
13 Apr 2026
Safety

Plasticity-Enhanced Multi-Agent Mixture of Experts for Dynamic Objective Adaptation in UAVs-Assisted Emergency Communication Networks

DGX agent

arXiv:2604.09028v1 Announce Type: cross Abstract: Unmanned aerial vehicles serving as aerial base stations can rapidly restore connectivity after disasters, yet abrupt changes in user mobility and tra

safetyarxiv-cs-lg
13 Apr 2026
Model Releases

SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment

DGX agent

arXiv:2604.08988v1 Announce Type: new Abstract: Current LLM-based agents demonstrate strong performance in episodic task execution but remain constrained by static toolsets and episodic amnesia, faili

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

SPASM: Stable Persona-driven Agent Simulation for Multi-turn Dialogue Generation

DGX agent

arXiv:2604.09212v1 Announce Type: new Abstract: Large language models are increasingly deployed in multi-turn settings such as tutoring, support, and counseling, where reliability depends on preservin

model-releasesarxiv-cs-cl
13 Apr 2026
Safety

Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection

DGX agent

arXiv:2604.07831v1 Announce Type: cross Abstract: Existing red-teaming studies on GUI agents have important limitations. Adversarial perturbations typically require white-box access, which is unavaila

safetyarxiv-cs-cl
10 Apr 2026
Agents

Exploring Plan Space through Conversation: An Agentic Framework for LLM-Mediated Explanations in Planning

DGX agent

arXiv:2603.02070v2 Announce Type: replace-cross Abstract: When automating plan generation for a real-world sequential decision problem, the goal is often not to replace the human planner, but to facil

agentsarxiv-cs-cl
10 Apr 2026
Safety

Learning to Negotiate: Multi-Agent Deliberation for Collective Value Alignment in LLMs

DGX agent

arXiv:2603.10476v2 Announce Type: replace Abstract: LLM alignment has progressed in single-agent settings through paradigms such as RL with human feedback (RLHF), while recent work explores scalable a

safetyarxiv-cs-cl
10 Apr 2026
Agents

Lighting-grounded Video Generation with Renderer-based Agent Reasoning

DGX agent

arXiv:2604.07966v1 Announce Type: new Abstract: Diffusion models have achieved remarkable progress in video generation, but their controllability remains a major limitation. Key scene factors such as

agentsarxiv-cs-cv
10 Apr 2026
Agents

AgonAlpha: Autonomous Alpha Discovery via Prompt Economy and Scalable Agentic Search

DGX agent

arXiv:2608.11250v1 Announce Type: new Abstract: Language models can propose many plausible trading factors, but an autonomous research system must also allocate its evaluation budget, verify its own e

agentsarxiv-cs-ai
13 Aug 2026
Model Releases

An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS

DGX agent

arXiv:2608.12249v1 Announce Type: new Abstract: Modernizing legacy Fortran is a problem of volume: the transformations are individually routine, but the codebases can be enormous, and across much of c

model-releasesarxiv-cs-ai
13 Aug 2026
Research

Backdoor Decontamination Dynamics in LLM Agents

DGX agent

arXiv:2608.11295v1 Announce Type: cross Abstract: Open-weight LLM agents are vulnerable to backdoors installed during fine-tuning, which may be undetectable if the trigger conditions are never met dur

researcharxiv-cs-ai
13 Aug 2026
Model Releases

FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents

DGX agent

arXiv:2608.11683v1 Announce Type: new Abstract: AI agents are increasingly deployed for professional investment research, yet no benchmark captures the complexity of the full investor workflow. Existi

model-releasesarxiv-cs-ai
13 Aug 2026
Model Releases

Harnessing agent memory to build lifelong AI partners for materials scientists

DGX agent

arXiv:2608.11224v1 Announce Type: new Abstract: Materials research advances through accumulated experience - scripts that work, protocols that are trusted, warnings attached to failed calculations or

model-releasesarxiv-cs-ai
13 Aug 2026
Safety

LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation

DGX agent

arXiv:2608.11967v1 Announce Type: cross Abstract: Large language model agents increasingly rely on long-horizon reasoning to solve complex tasks involving planning, tool use, and memory. A critical ca

safetyarxiv-cs-ai
13 Aug 2026
Safety

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL

DGX agent

arXiv:2608.12253v1 Announce Type: cross Abstract: Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We show that

safetyarxiv-cs-ai
13 Aug 2026
Hardware

Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control

DGX agent

arXiv:2608.12123v1 Announce Type: cross Abstract: LLM-agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome, update state, and emit the next

hardwarearxiv-cs-ai
13 Aug 2026
Model Releases

RecSys Factory: Bounding LLM Agent Autonomy to Decision Points in the Industrial Recommender Lifecycle

DGX agent

arXiv:2608.11241v1 Announce Type: new Abstract: Deploying LLM agents into industrial recommender operations exposes a three-way tension we frame as the autonomy-determinism-efficiency trilemma: genera

model-releasesarxiv-cs-ai
13 Aug 2026
Model Releases

360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents

DGX agent

arXiv:2608.08814v1 Announce Type: cross Abstract: We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environment construc

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography

DGX agent

arXiv:2608.07651v1 Announce Type: new Abstract: Large language models (LLMs) show promise in medical image interpretation but suffer from hallucination, limited accuracy, and run-to-run inconsistency.

model-releasesarxiv-cs-ai
11 Aug 2026
Safety

An Agentic Generative Large Language Model for Treatment Planning of Colorectal Cancer

DGX agent

arXiv:2608.09142v1 Announce Type: new Abstract: Treatment planning in precision oncology requires synthesizing heterogeneous patient information with rapidly evolving clinical guidelines to ensure gui

safetyarxiv-cs-cl
11 Aug 2026
Model Releases

BibTeX Citation Errors in Scientific Publishing Agents: Evaluation and Mitigation

DGX agent

arXiv:2604.03159v2 Announce Type: replace-cross Abstract: Large language models with web search are increasingly used in scientific publishing agents, yet they produce BibTeX entries with pervasive fi

model-releasesarxiv-cs-cl
11 Aug 2026
Agents

CARD: Controlled Agentic Reddit Discussions for Credit Card Simulation

DGX agent

arXiv:2608.09790v1 Announce Type: new Abstract: Online credit card discussions provide a natural setting for studying how consumers communicate about financial products. Simulating these discussions r

agentsarxiv-cs-ai
11 Aug 2026
Model Releases

ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with Coupons

DGX agent

arXiv:2608.09282v1 Announce Type: new Abstract: Real-world shopping often requires constructing a basket of complementary items rather than retrieving a single product. Such combo-shopping tasks arise

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses

DGX agent

arXiv:2608.08466v1 Announce Type: new Abstract: Modern LLM agents are often improved by modifying prompts, tools, or workflows manually, while the executable scaffold surrounding the model---the harne

model-releasesarxiv-cs-ai
11 Aug 2026
Safety

Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards

DGX agent

arXiv:2608.07531v1 Announce Type: cross Abstract: Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence. Existing ext

safetyarxiv-cs-ai
11 Aug 2026
Model Releases

Towards Researcher Agents for Knowledge-Graph Question Answering

DGX agent

arXiv:2608.07700v1 Announce Type: new Abstract: Translating a natural-language question into a SPARQL query that can be executed against a large knowledge graph requires resolving lexical ambiguity, g

model-releasesarxiv-cs-ai
11 Aug 2026
Research

Tree-of-Experience: Hierarchical Experience Management for Self-Evolving Agents

DGX agent

arXiv:2608.09044v1 Announce Type: new Abstract: Continual self-evolution requires LLM agents to transform environmental interactions into reliable and reusable experience. Existing methods typically r

researcharxiv-cs-cl
11 Aug 2026
Safety

What Keeps Agent Skills from Being Reusable? Evidence from 138K SKILL.md Files

DGX agent

arXiv:2608.08453v1 Announce Type: new Abstract: Under the current standard, Agent Skills are SKILL.md files that combine instructions with supporting resources, enabling Large Language Model (LLM) age

safetyarxiv-cs-ai
11 Aug 2026
Agents

Explicit, Not Longer: What Makes Epistemic Stance Survive Memory Compression

DGX agent

arXiv:2608.06953v1 Announce Type: cross Abstract: Agent memory systems compress what they store, and compression is built to drop qualifiers, so a claim's epistemic standing tends not to survive being

agentsarxiv-cs-ai
10 Aug 2026
Model Releases

MemWM: Memory-Augmented Text-Based World Model

DGX agent

arXiv:2608.07107v1 Announce Type: new Abstract: World models are increasingly used to support planning in agents by predicting how environment states evolve in response to agent actions. Yet fluent ne

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

A Two-Tier Perspective on Inference-Time Parallelism in Multi-Agent LLM Systems

DGX agent

arXiv:2608.05791v1 Announce Type: cross Abstract: Large language model (LLM)-driven multi-agent systems typically require multiple model invocations and complex coordination during inference, and thei

model-releasesarxiv-cs-ai
7 Aug 2026
← Previous
1…8485868788…236
Next →