AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

84,606Total entries
1Added by human
84,605Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,973 results
1 Jun 2026

Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation

Model ReleasesDGX agent

arXiv:2605.30916v1 Announce Type: new Abstract: AI benchmarks have well-documented limitations, with prior work examining contamination, saturation, and construct underspecification. Aggregation has r

29 May 2026

Analyzing Persona Effects in Generated Explanations from Multimodal LLM Agents in Urban Perception

ResearchDGX agent

arXiv:2605.29064v1 Announce Type: new Abstract: We study how persona prompting shapes language generated by multimodal large language models in an urban perception setting. Using 59,808 annotations fr

AutoSizer: Automatic Sizing of Analog and Mixed-Signal Circuits via Large Language Model (LLM) Agents

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2602.02849v2 Announce Type: replace Abstract: The design of Analog and Mixed-Signal (AMS) integrated circuits remains heavily reliant on expert knowledge, with transistor sizing a major bottlene

DirectorBench: Diagnosing Long-Form Video Generation with Personalized Multi-Agent Evaluation

Model ReleasesDGX agent

arXiv:2605.30090v1 Announce Type: new Abstract: Long-form video generation is rapidly moving from short, single-scene synthesis toward minute-long, multi-shot creation with narrative structure, cinema

Harmonizing Real-Time Constraints and Long-Horizon Reasoning: An Asynchronous Agentic Framework for Dynamic Scheduling

SafetyDGX agent

arXiv:2605.29262v1 Announce Type: new Abstract: The Dynamic Flexible Job Shop Scheduling Problem (DFJSP) necessitates a trade-off between instant reaction to stochastic disturbances and global optimiz

Learning Design Skills as Memory Policies for Agentic Photonic Inverse Design

Model ReleasesDGX agent

arXiv:2605.29421v1 Announce Type: new Abstract: Photonic crystal fiber (PCF) inverse design remains challenging because candidate geometries must satisfy coupled optical targets under expensive electr

Loong: A Human-Like Long Document Translation Agent with Observe-and-Act Adaptive Context Selection

Model ReleasesDGX agent

arXiv:2605.30274v1 Announce Type: cross Abstract: Document-level translation remains one of the most challenging tasks for large language models, which are constrained by limited context windows that

Physics Is All You Need? A Case Study in Physicist-Supervised AI Development of Scientific Software

Model ReleasesDGX agent

arXiv:2605.30353v1 Announce Type: new Abstract: Are AI agents tools, co-authors, or researchers? We present a quantified case study (N=1): a physicist supervising an AI coding agent (Claude Code, Sonn

Small Agent Group is the Future of Digital Health

Model ReleasesDGX agent

arXiv:2602.08013v2 Announce Type: replace Abstract: The rapid adoption of large language models (LLMs) in digital health has been driven by a 'scaling-first' philosophy, i.e., the assumption that clin

VitalAgent: A Tool-Augmented Agent for Reactive and Proactive Physiological Monitoring over Wearable Health Data

Model ReleasesDGX agent

arXiv:2605.29483v1 Announce Type: new Abstract: Wearable devices enable continuous monitoring of physiological signals such as ECG and PPG, but existing mHealth systems are largely limited to task-spe

28 May 2026

CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning

Model ReleasesDGX agent

arXiv:2605.28056v1 Announce Type: new Abstract: Portrait animation methods have achieved substantial visual quality and lip synchronization, but fine-grained manipulation of the eye region still faces

CyberJurors: A Multi-Agent Simulation Task for E-Commerce Disputes Verdict

Model ReleasesDGX agent

arXiv:2605.28369v1 Announce Type: new Abstract: E-commerce platforms have begun recruiting crowdsourced jurors to adjudicate massive volumes of transaction disputes. Unlike formal legal judgment, E-co

DSSE: a drone swarm search environment

AgentsDGX agent

arXiv:2307.06240v2 Announce Type: replace-cross Abstract: The Drone Swarm Search project is an environment, based on extsc{PettingZoo}, that is to be used in conjunction with multi-agent (or single-ag

KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs

Model ReleasesDGX agent

arXiv:2605.27984v1 Announce Type: cross Abstract: Speech language models (SpeechLMs) have achieved substantial progress by extending large language models (LLMs) to the speech modality. However, Speec

MangaFlow: An End-to-End Agentic Framework for Controllable Story to Manga Generation

Model ReleasesDGX agent

arXiv:2605.28173v1 Announce Type: new Abstract: End-to-end manga generation is a structured visual storytelling task that requires story decomposition, recurring character and scene grounding, page la

MGRetrieval: Memory-Guided Reflective Retrieval for Long-Term Dialogue Agents

ResearchDGX agent

arXiv:2605.27437v1 Announce Type: cross Abstract: Large Language Models (LLMs) have made significant progress in dialogue, yet redundant memory contexts severely limit their effectiveness in long-term

RAG-Coding: Enhancing LLM Medical Coding with Structured External Knowledge

AgentsDGX agent

arXiv:2605.27377v1 Announce Type: cross Abstract: We present RAG-Coding, an agentic method for automated ICD-10-CM coding. RAG-Coding orchestrates four large language model (LLM) agents and grounds th

Reflective Dialogue between Teacher and Solver Agents for Video Question Answering

Model ReleasesDGX agent

arXiv:2605.27885v1 Announce Type: new Abstract: Various approaches have been proposed to adapt Vision-Language Models (VLMs) to specialized domains for Video Question Answering, including fine-tuning

The Script is All You Need: An Agentic Framework for Long-Horizon Dialogue-to-Cinematic Video Generation

Model ReleasesDGX agent

arXiv:2601.17737v3 Announce Type: replace-cross Abstract: Recent advances in video generation have produced models capable of synthesizing stunning visual content from simple text prompts. However, th

Why LLMs Fail at Causal Discovery and How Interventional Agents Escape

Model ReleasesDGX agent

arXiv:2605.27567v1 Announce Type: new Abstract: Causal discovery is a cornerstone of scientific reasoning, yet whether large language models can perform it reliably remains an open question. Recent be

27 May 2026

BEAT: Rhythm-Elastic Alignment for Agentic Music-guided Movie Trailer Generation

Model ReleasesDGX agent

arXiv:2605.27067v1 Announce Type: new Abstract: Automatic movie trailer generation must select shots from a full-length film and synchronize them with background music. Existing methods either relegat

Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2605.26684v1 Announce Type: cross Abstract: Group-based reinforcement learning (RL) methods have achieved remarkable success in improving the performance of large language models (LLMs) and have

From production traces to better AI agents: Automating the LLMOps feedback loop

ApplicationsDGX agent

Production AI traces are the raw material for better evals, prompts, datasets, and fine-tuned models. This post shows how the Arize AX Airflow Provider turns that feedback loop into scheduled, monitor

Grammar of the Wave: Towards Explainable Multivariate Time Series Event Detection via Neuro-Symbolic VLM Agents

Model ReleasesDGX agent

arXiv:2603.11479v2 Announce Type: replace-cross Abstract: Time Series Event Detection (TSED) aims to localize semantically meaningful events in time series data, with critical applications in high-sta

insanely good company to keep

AgentsDGX agent

insanely good company to keep 🆕Railway's Agent-Native Cloud: 3M users, 100K signups/week, $200K+ coding agent spend, production forks, & the death of PRs https://latent.space/p/railway @Railway founde

Maat: The Agentic Legal Research Assistant for Competition Protection

Model ReleasesDGX agent

arXiv:2605.27331v1 Announce Type: new Abstract: Competition law experts conducting legal research must review extensive volumes of cases, decisions, and judicial reports to identify precedents and ass

REVERSE: Reinforcing Evidence Verification and Search for Agentic Image geo-localization

Local AiDGX agent

arXiv:2605.26861v1 Announce Type: new Abstract: Image geo-localization aims to determine where a photograph was taken, a task that often requires more than recognizing visible landmarks. Human experts

Uncertainty-Aware Gaussian Map for Vision-Language Navigation

AgentsDGX agent

arXiv:2605.26503v1 Announce Type: new Abstract: Vision-Language Navigation (VLN) requires an agent to navigate 3D environments following natural language instructions. During navigation, existing agen

Your agents can burn through $10k overnight before you notice. LangSmith LLM Gateway stops that. The platform where you already observe, eva…

ApplicationsDGX agent

LangSmith's LLM Gateway is a cost control and monitoring tool designed to prevent unexpected high spending on LLM API calls by providing observability and likely rate-limiting or budget-cap features.

26 May 2026

CRPO: Character-centric Group Relative Policy Optimization for Role-aware Reasoning in Role-playing Agents

SafetyDGX agent

arXiv:2605.25511v1 Announce Type: new Abstract: Recent advancements in Reinforcement Learning (RL), particularly Group Relative Policy Optimization (GRPO), have significantly enhanced the reasoning ca

Learning to Trust: Bayesian Adaptation to Varying Suggester Reliability in Sequential Decision Making

AgentsDGX agent

arXiv:2511.12378v2 Announce Type: replace Abstract: Autonomous agents operating in sequential decision-making tasks under uncertainty can benefit from external action suggestions, which provide valuab

Millions of AI agents imperiled by critical vulnerability in open source package

IndustryDGX agent

A critical vulnerability (CVE-2026-33579, scoring 9.8/10 severity) was discovered in the OpenClaw open source package, allowing attackers with minimal access to silently escalate privileges to full ad

Prior Policy Guided Dual-Agent Coordinated Manipulation Planning of Spacecraft-Manipulator System

SafetyDGX agent

arXiv:2605.25362v1 Announce Type: new Abstract: The strong dynamic coupling between the manipulator and the base poses a significant challenge to maintaining spacecraft attitude stability, potentially

25 May 2026

CHRONOS: Temporally-Aware Multi-Agent Coordination for Evolving Data Marketplaces

Model ReleasesDGX agent

arXiv:2605.23887v1 Announce Type: cross Abstract: Temporal knowledge-graph data marketplaces face three coupled failures in static designs: stale hybrid index shortcuts reduce recall as edges evolve,

Human-in-the-Loop Multi-Agent Ventilator Decision Support with Contextual Bandit Preference Learning

SafetyDGX agent

arXiv:2605.23320v1 Announce Type: new Abstract: Ventilator decision support requires sequential decisions that track evolving physiology and disease trajectories while respecting safety boundaries and

24 May 2026

Built an AI screen memory using llama.cpp + Gemma 4 — remembers everything you do on your computer,search/chat or make agents over it. 100% local

Model ReleasesDGX agent

This project demonstrates a local AI system built with llama.cpp and Gemma 4 that captures and analyzes screen activity to create persistent memory of user computer interactions, enabling search, chat

23 May 2026

Cross-domain benchmarks reveal when coordinated AI agents improve scientific inference from partial evidence

Model ReleasesDGX agent

arXiv:2605.22300v1 Announce Type: cross Abstract: Scientific evidence often spans instruments, databases, and disciplines, so no single source records the full phenomenon. This makes it difficult to d

22 May 2026

Generic LLMs won’t cut it — AI agents demand unified context

IndustryDGX agent

Unified AI service management lives or dies on context — and that context demands consolidating fragmented tools, data and operations into a single foundation. That consolidation imperative is now the

GitHub recognized as a Leader in the Gartner® Magic Quadrant™ for Enterprise AI Coding Agents for the third year in a row

ApplicationsDGX agent

We are committed to empowering every developer by building an open, secure, and AI-powered platform that defines the future of software development. The post GitHub recognized as a Leader in the Gartn

Matching with Deliberation: Test-Time Evolutionary Hierarchical Multi-Agents for Zero-Shot Compositional Image Retrieval

Model ReleasesDGX agent

arXiv:2605.22478v1 Announce Type: new Abstract: Zero-Shot Compositional Image Retrieval (ZS-CIR) requires both preserving the visual continuity of the reference image and faithfully executing the sema

Mega AI IPOs incoming, Google’s agentic blitz and Nvidia’s next big business

HardwareDGX agent

Now we know for sure: This will be the year of monster initial public offerings. Elon Musk’s SpaceX filed for an IPO this week, aiming to raise a record $80 billion or more, and OpenAI was expected to

Quality and Security Signals in AI-Generated Python Refactoring Pull Requests

AgentsDGX agent

arXiv:2605.21453v1 Announce Type: cross Abstract: As AI agents increasingly contribute to code development and maintenance, there is still limited empirical evidence on the quality and risk characteri

recommended reading.

AgentsDGX agent

recommended reading. babyagi has ~200 citations, but 0 papers... i just published my first paper on arXiv 😆 'The Log is the Agent: Event-Sourced Reactive Graphs for Auditable, Forkable Agentic Systems

21 May 2026

Agentic Physical AI toward a Domain-Specific Foundation Model for Nuclear Reactor Control

Model ReleasesDGX agent

arXiv:2512.23292v3 Announce Type: replace-cross Abstract: The prevailing paradigm in AI for physical systems (scaling general-purpose foundation models toward universal multimodal reasoning) confronts

Mem-pi: Adaptive Memory through Learning When and What to Generate

AgentsDGX agent

arXiv:2605.21463v1 Announce Type: new Abstract: We present Mem-pi, a framework for adaptive memory in large language model (LLM) agents, where useful guidance is generated on demand rather than retrie

Time-To-Reach Separation and Safety Filtering for Safe, Fair, and Efficient Multi-Agent Coordination

SafetyDGX agent

arXiv:2605.20625v1 Announce Type: cross Abstract: Advanced Air Mobility (AAM) operations are expected to significantly increase aerial traffic in urban airspace, requiring autonomous traffic managemen

20 May 2026

A Case for Agentic Tuning: From Documentation to Action in PostgreSQL

Model ReleasesDGX agent

arXiv:2605.19988v1 Announce Type: cross Abstract: Documentation has long guided computer system tuning by distilling expert knowledge into per-parameter recommendations. Yet such guides capture only w

CopT: Contrastive On-Policy Thinking with Continuous Spaces for General and Agentic Reasoning

SafetyDGX agent

arXiv:2605.20075v1 Announce Type: cross Abstract: Chain-of-thought (CoT) is a standard approach for eliciting reasoning capabilities from large language models (LLMs). However, the common CoT paradigm

From Prompts to Pavement Through Time: Temporal Grounding in Agentic Scene-to-Plan Reasoning

Model ReleasesDGX agent

arXiv:2605.19824v1 Announce Type: new Abstract: Recent attempts to support high-level scene interpretation and planning in Autonomous Vehicles (AVs) using ensembles of Large Language Models (LLMs) and

RLFTSim: Realistic and Controllable Multi-Agent Traffic Simulation via Reinforcement Learning Fine-Tuning

SafetyDGX agent

arXiv:2605.19033v1 Announce Type: cross Abstract: Supervised open-loop training has been widely adopted for training traffic simulation models; however, it fails to capture the inherently dynamic, mul

What we learned testing 7 models under the same agent harness

Model ReleasesDGX agent

Model swaps look like configuration changes, but they behave more like product migrations. A new model may be cheaper, faster, easier to get capacity for, or stronger on public benchmarks.... The post

19 May 2026

Causal Influences over Social Learning Networks

AgentsDGX agent

arXiv:2307.09575v2 Announce Type: replace-cross Abstract: This paper investigates causal influences between agents linked by a social graph and interacting over time. In particular, the work examines

Diagnosing Korean-Language LLM Political Bias via Census-Grounded Agent Simulation

Local AiDGX agent

arXiv:2605.18395v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit systematic political biases in voter simulations, but their underlying mechanisms and cross-lingual generalizatio

Episodic-Semantic Memory Architecture for Long-Horizon Scientific Agents

Model ReleasesDGX agent

arXiv:2605.17625v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into persistent scientific collaborators, context window saturation has emerged as a critical bottleneck. Scienti

Goal inference with Rao-Blackwellized Particle Filters

AgentsDGX agent

arXiv:2512.09269v2 Announce Type: replace Abstract: Inferring the eventual goal of a mobile agent from noisy observations of its trajectory is a fundamental estimation problem. We initiate the study o

Google just redesigned the search box for the first time in 25 years — here’s why it matters more than you think.

Model ReleasesDGX agent

For a quarter century, the Google search box has been one of the most recognizable interfaces in computing: a thin white rectangle, a blinking cursor, a few typed words, and a list of blue links. On T

Google reimagines search with AI agents and generative interfaces

IndustryDGX agent

Google LLC’s search is getting an artificial intelligence makeover today, as the company announced during its Google I/O developer conference that it’s upgrading the search experience to “reimagine it

Google’s new Gemini 3.5 Flash is the clear leader on the Intelligence vs Speed Pareto frontier and makes large gains on GDPval-AA (real-worl…

Model ReleasesDGX agent

Google’s new Gemini 3.5 Flash is the clear leader on the Intelligence vs Speed Pareto frontier and makes large gains on GDPval-AA (real-world agentic tasks), but is 5x the cost of Gemini 3 Flash @Goog

GVGAI-LLM: Evaluating Large Language Model Agents with Infinite Games

Model ReleasesDGX agent

arXiv:2508.08501v3 Announce Type: replace Abstract: We introduce GVGAI-LLM, a video game benchmark for evaluating the reasoning and problem-solving capabilities of large language models (LLMs). Built

MirrorBench: A Benchmark to Evaluate Conversational User-Proxy Agents for Human-Likeness

Model ReleasesDGX agent

arXiv:2601.08118v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used as human simulators, both for evaluating conversational systems and for generating fine-tuning da

← Previous
1…146147148149150…300
Next →