AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,289 results
Model Releases

Helicase: Uncertainty-Guided Supply Chain Knowledge Graph Construction with Autonomous Multi-Agent LLMs

DGX agent

arXiv:2605.26835v1 Announce Type: new Abstract: LLM-based multi-agent systems have been widely adopted for knowledge retrieval and report generation, synthesizing known information through web search

model-releasesarxiv-cs-ai
27 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Agents

Probing the Knowledge Boundary: An Interactive Agentic Framework for Deep Knowledge Extraction

DGX agent

arXiv:2602.00959v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) can be seen as compressed knowledge bases, but it remains unclear what knowledge they truly contain and how far t

agentsarxiv-cs-cl
27 May 2026
Research

RepoMirage: Probing Repository Context Reasoning in Code Agents with Perturbations

DGX agent

arXiv:2605.26177v1 Announce Type: cross Abstract: Code agents are currently having skillful performance on repository-level software engineering benchmarks, but it remains unclear whether success on e

researcharxiv-cs-ai
27 May 2026
Safety

StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning

DGX agent

arXiv:2605.27140v1 Announce Type: new Abstract: Reinforcement learning for multi-turn agents suffers from a credit-assignment mismatch: rewards are sparse and trajectory-level, while success often hin

safetyarxiv-cs-ai
27 May 2026
Agents

A Multi-Agent LLM Framework for Rating the Quality of Surgical Feedback

DGX agent

arXiv:2605.25440v1 Announce Type: cross Abstract: Verbal feedback delivered by attending surgeons in the operating room plays a critical formative role in resident trainee skill acquisition. Yet, asse

agentsarxiv-cs-ai
26 May 2026
Model Releases

AvalancheBench: Evaluating Enterprise Data Agents Through Latent World Recovery

DGX agent

arXiv:2605.24183v1 Announce Type: cross Abstract: We introduce AvalancheBench, a benchmark for evaluating enterprise data agents through latent world recovery. AvalancheBench improves on existing benc

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Can LLMs Time Travel? Enhancing Temporal Consistency in Legal Agentic Search through Reinforcement Learning

DGX agent

arXiv:2605.25920v1 Announce Type: cross Abstract: While large language models (LLMs) augmented with agentic search capabilities show promise for legal reasoning, they overlook a fundamental constraint

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents

DGX agent

arXiv:2605.25624v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has driven breakthroughs in domains such as math, tool-use, and software engineering, yet its exte

model-releasesarxiv-cs-ai
26 May 2026
Safety

Hide-and-Shill: A Reinforcement Learning Framework for Market Manipulation Detection in Symphony-a Decentralized Multi-Agent System

DGX agent

arXiv:2507.09179v3 Announce Type: replace Abstract: Decentralized finance (DeFi) has introduced a new era of permissionless financial innovation but also led to unprecedented market manipulation. With

safetyarxiv-cs-ai
26 May 2026
Model Releases

Insuring Every Action: An Authority Frontier Framework for Runtime Actuarial Control of Autonomous AI Agents

DGX agent

arXiv:2605.25632v1 Announce Type: new Abstract: Autonomous AI agents increasingly issue side-effect-bearing actions: database mutations, refunds, payments, external commitments. We propose the Actuari

model-releasesarxiv-cs-ai
26 May 2026
Safety

MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents

DGX agent

arXiv:2602.02474v2 Announce Type: replace-cross Abstract: Most Large Language Model (LLM) agent memory systems rely on a small set of static, hand-designed operations for extracting memory. These fixe

safetyarxiv-cs-ai
26 May 2026
Model Releases

MimirRAG: A Multi-Agent RAG Framework for Financial Data Retrieval with Metadata Integration

DGX agent

arXiv:2605.25030v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) systems offer a promising approach to reduce hallucinations and improve answer accuracy in large language models (L

model-releasesarxiv-cs-lg
26 May 2026
Safety

PolyGnosis 2.0: Enhancing LLM Reasoning via Agentic Harness Engineering for Polymarket and OSINT Insight Extraction

DGX agent

arXiv:2605.25958v1 Announce Type: new Abstract: This paper introduces PolyGnosis 2.0, a pioneering multi-agent architecture designed to extract predictive intelligence by synthesizing Polymarket anoma

safetyarxiv-cs-cl
26 May 2026
Model Releases

SODE: Analyzing Social Dynamics in LLM Agents

DGX agent

arXiv:2605.23949v1 Announce Type: cross Abstract: As Large Language Models (LLMs) evolve into interactive agents, understanding their behavioral alignment within human social dynamics becomes essentia

model-releasesarxiv-cs-ai
26 May 2026
Safety

Stop Comparing LLM Agents Without Disclosing the Harness

DGX agent

arXiv:2605.23950v1 Announce Type: new Abstract: This position paper argues that, for long-horizon tasks evaluated across models with comparable frontier capability, the agent execution harness, namely

safetyarxiv-cs-ai
26 May 2026
Model Releases

Understanding Conversational Patterns in Multi-agent Programming: A Case Study on Fibonacci Game Development

DGX agent

arXiv:2605.24138v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly applied to software engineering (SE), yet their potential for autonomous, role-oriented collaboration re

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Agentic-VLA: Efficient Online Adaptation for Vision-Language-Action Models

DGX agent

arXiv:2605.22896v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising paradigm for robotic manipulation by leveraging pre-trained vision-language representa

model-releasesarxiv-cs-ai
25 May 2026
Safety

Energy per Successful Goal: Goal-Level Energy Accounting for Agentic AI Systems

DGX agent

arXiv:2605.22883v1 Announce Type: new Abstract: Current AI energy benchmarks measure consumption at the granularity of a single model invocation or training run. For classical single-turn workloads th

safetyarxiv-cs-ai
25 May 2026
Model Releases

Evaluating Memory Structure in LLM Agents

DGX agent

arXiv:2602.11243v2 Announce Type: replace-cross Abstract: Modern LLM-based agents and chat assistants rely on long-term memory frameworks to store reusable knowledge, recall user preferences, and augm

model-releasesarxiv-cs-cl
25 May 2026
Safety

IntentScore: Intent-Conditioned Action Evaluation for Computer-Use Agents

DGX agent

arXiv:2604.05157v2 Announce Type: replace Abstract: Computer-Use Agents (CUAs) leverage large language models to execute GUI operations on desktop environments, yet they generate actions without evalu

safetyarxiv-cs-ai
25 May 2026
Agents

KPI2KVI: A Multi Agent Workflow for Calculating Key Value Indicators from Service Descriptions

DGX agent

arXiv:2605.22825v1 Announce Type: cross Abstract: Key Value Indicators (KVIs) provide a decision oriented view of a service by summarizing how operational performance translates into stakeholder value

agentsarxiv-cs-ai
25 May 2026
Agents

NeuroWeaver: An Autonomous Evolutionary Agent for Exploring the Programmatic Space of EEG Analysis Pipelines

DGX agent

arXiv:2602.13473v2 Announce Type: replace Abstract: Although foundation models have demonstrated remarkable success in general domains, the application of these models to electroencephalography (EEG)

agentsarxiv-cs-ai
25 May 2026
Model Releases

PrefBench: Evaluating Zero-Shot LLM Agents in Hidden-Preference Personalized Pricing Negotiations

DGX agent

arXiv:2605.22855v1 Announce Type: cross Abstract: Personalized pricing negotiations are a challenging testbed for LLM agents because successful interaction does not guarantee profitable decision makin

model-releasesarxiv-cs-ai
25 May 2026
Safety

SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety

DGX agent

arXiv:2605.05704v2 Announce Type: replace-cross Abstract: Recent advances in foundation models have transformed LLMs from passive conversational systems into autonomous agents capable of reasoning and

safetyarxiv-cs-ai
25 May 2026
Research

SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents

DGX agent

arXiv:2602.11210v4 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become a key paradigm for training software engineering (SWE) agents, but existing pipelines typically rely on

researcharxiv-cs-lg
23 May 2026
Safety

Governance by Design: Architecting Agentic AI for Organizational Learning and Scalable Autonomy

DGX agent

arXiv:2605.20210v1 Announce Type: cross Abstract: Agentic AI systems - systems that can pursue goals through multi-step planning and tool-mediated action with limited direct supervision - are moving f

safetyarxiv-cs-ai
22 May 2026
Agents

ImProver: Agent-Based Automated Proof Optimization

DGX agent

arXiv:2410.04753v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have been used to generate formal proofs of mathematical theorems in proofs assistants such as Lean. However, we

agentsarxiv-cs-cl
22 May 2026
Model Releases

Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents

DGX agent

arXiv:2605.22148v1 Announce Type: cross Abstract: Self-evolving skill libraries, pioneered by Voyager, let frozen LLM agents accumulate reusable knowledge without weight updates, yet recent evaluation

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation

DGX agent

arXiv:2605.20204v1 Announce Type: cross Abstract: LLM-based user simulation is the primary mechanism for end-to-end agent evaluation, yet simulated users are poor proxies for real humans: unconstraine

model-releasesarxiv-cs-ai
22 May 2026
Agents

STRUCTSENSE: A Task-Agnostic Agentic Framework for Structured Information Extraction with Human-In-The-Loop Evaluation and Benchmarking

DGX agent

arXiv:2507.03674v3 Announce Type: replace Abstract: Extracting structured information from scientific literature is critical for accelerating discovery, yet Large Language Models (LLMs) often struggle

agentsarxiv-cs-cl
22 May 2026
Safety

X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction

DGX agent

arXiv:2605.05765v2 Announce Type: replace Abstract: Inspired by the development of OpenClaw, there is a growing demand for mobile-based personal agents capable of handling complex and intuitive intera

safetyarxiv-cs-cv
22 May 2026
Model Releases

AgentAtlas: Beyond Outcome Leaderboards for LLM Agents

DGX agent

arXiv:2605.20530v1 Announce Type: cross Abstract: Large language model agents now act on codebases, browsers, operating systems, calendars, files, and tool ecosystems, but the benchmarks used to evalu

model-releasesarxiv-cs-cl
21 May 2026
Model Releases

Evolutionary Generation of Multi-Agent Systems

DGX agent

arXiv:2602.06511v3 Announce Type: replace Abstract: Large language model (LLM)-based multi-agent systems (MAS) show strong promise for complex reasoning, planning, and tool-augmented tasks, but design

model-releasesarxiv-cs-lg
21 May 2026
Safety

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents

DGX agent

arXiv:2605.20246v1 Announce Type: new Abstract: Recently, vision-language model (VLM) agents have shown promising progress in open-world tasks, where successful task completion often requires multiple

safetyarxiv-cs-lg
21 May 2026
Model Releases

Training Language Agents to Learn from Experience

DGX agent

arXiv:2605.20477v1 Announce Type: cross Abstract: Language agents can adapt from experience in interactive environments, but current reflection-based methods can only self-correct within a single task

model-releasesarxiv-cs-cl
21 May 2026
Model Releases

WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents

DGX agent

arXiv:2605.20306v1 Announce Type: new Abstract: We introduce WildRoadBench, a wild aerial road-damage grounding benchmark that couples direct visual grounding by vision-language models with autonomous

model-releasesarxiv-cs-cv
21 May 2026
Agents

A Closed-loop, State-centric, Multi-agent Framework for Passenger Load Estimation from Heterogeneous Data Streams

DGX agent

arXiv:2605.19834v1 Announce Type: cross Abstract: To support operations and passenger-facing services, transit agencies need reliable passenger load trajectories. Currently, load estimates are typical

agentsarxiv-cs-ai
20 May 2026
Agents

A novel YOLO26-MoE optimized by an LLM agent for insulator fault detection considering UAV images

DGX agent

arXiv:2605.19595v1 Announce Type: cross Abstract: The inspection of electrical power line insulators is essential for ensuring grid reliability and preventing failures caused by damaged or degraded in

agentsarxiv-cs-ai
20 May 2026
Agents

Agentic Discovery of Cryomicroneedle Formulations

DGX agent

arXiv:2605.19677v1 Announce Type: new Abstract: Cryomicroneedles offer a route to minimally invasive intradermal delivery of living cells, but their cryogenic formulations must reconcile cell protecti

agentsarxiv-cs-lg
20 May 2026
Agents

CMAD: Cooperative Multi-Agent Diffusion via Stochastic Optimal Control

DGX agent

arXiv:2602.10933v2 Announce Type: replace Abstract: Continuous-time generative models have achieved remarkable success in image restoration and synthesis. However, controlling the composition of multi

agentsarxiv-cs-lg
20 May 2026
Model Releases

CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing

DGX agent

arXiv:2605.19484v1 Announce Type: cross Abstract: While GUI agents have made significant progress in web navigation and basic operating system tasks, their capabilities in professional creative workfl

model-releasesarxiv-cs-ai
20 May 2026
Safety

Governing Evolving Memory in LLM Agents: Risks, Mechanisms, and the Stability and Safety Governed Memory (SSGM) Framework

DGX agent

arXiv:2603.11768v2 Announce Type: replace Abstract: Long-term memory has emerged as a foundational component of autonomous Large Language Model (LLM) agents, enabling continuous adaptation, lifelong m

safetyarxiv-cs-ai
20 May 2026
Model Releases

Hallucination as Exploit: Evidence-Carrying Multimodal Agents

DGX agent

arXiv:2605.19192v1 Announce Type: new Abstract: Multimodal agents use screenshots, documents, and webpages to choose tool calls. When a false visual claim triggers a click, email, extraction, or trans

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

POLAR-Bench: A Diagnostic Benchmark for Privacy-Utility Trade-offs in LLM Agents

DGX agent

arXiv:2605.19127v1 Announce Type: new Abstract: LLM agents increasingly have access to private user data and act on the user's behalf when interacting with third-party systems. The user defines what m

model-releasesarxiv-cs-ai
20 May 2026
Agents

Towards Discovery of Polymers for Insulin Delivery via Physics-Grounded Agentic Workflows

DGX agent

arXiv:2605.18831v1 Announce Type: cross Abstract: Cold-chain storage limits access to insulin for hundreds of millions of people; a thermally protective patch polymer could help, but the design space

agentsarxiv-cs-lg
20 May 2026
Agents

Agentic Cost-Aware Query Planning with Knowledge Distillation for Big Data Analytics

DGX agent

arXiv:2605.17831v1 Announce Type: new Abstract: Query optimization in big data analytics remains computationally expensive, particularly for resource-constrained environments where traditional optimiz

agentsarxiv-cs-lg
19 May 2026
Model Releases

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech

DGX agent

arXiv:2605.17583v1 Announce Type: new Abstract: While existing text-to-speech (TTS) models exhibit high expressiveness, fine-grained control over composite instructions remains challenging due to the

model-releasesarxiv-cs-cv
19 May 2026
Research

AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents

DGX agent

arXiv:2605.17933v1 Announce Type: new Abstract: Vision-language model (VLM) agents increasingly rely on memory-augmented reinforcement learning to reuse experience across long-horizon tasks, yet most

researcharxiv-cs-cv
19 May 2026
← Previous
1…8081828384…236
Next →