AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
15 May 2026

Prompting Policies for Multi-step Reasoning and Tool-Use in Black-box LLMs with Iterative Distillation of Experience

SafetyDGX agent

arXiv:2605.14443v1 Announce Type: new Abstract: The shift toward interacting with frozen, 'black-box' Large Language Models (LLMs) has transformed prompt engineering from a heuristic exercise into a c

ProtoMedAgent: Multimodal Clinical Interpretability via Privacy-Aware Agentic Workflows

AgentsDGX agent

arXiv:2605.14113v1 Announce Type: cross Abstract: While interpretable prototype networks offer compelling case-based reasoning for clinical diagnostics, their raw continuous outputs lack the semantic

PROVE: A Perceptual RemOVal cohErence Benchmark for Visual Media

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.14534v1 Announce Type: cross Abstract: Evaluating object removal in images and videos remains challenging because the task is inherently one-to-many, yet existing metrics frequently disagre

Proximal Action Replacement for Behavior Cloning Actor-Critic in Offline Reinforcement Learning

SafetyDGX agent

arXiv:2602.07441v2 Announce Type: replace-cross Abstract: Offline reinforcement learning (RL), which optimizes policies using a previously collected static dataset, is an important branch of RL. A pop

PyCSP3-Scheduling: A Scheduling Extension for PyCSP3

ResearchDGX agent

arXiv:2605.14559v1 Announce Type: new Abstract: PyCSP^3 provides a productive way to build constraint models for solving combinatorial constrained problems and export them to XCSP^3, preserving a comp

Quantifying and Mitigating Premature Closure in Frontier LLMs

SafetyDGX agent

arXiv:2605.15000v1 Announce Type: cross Abstract: Premature closure, or committing to a conclusion before sufficient information is available, is a recognized contributor to diagnostic error but remai

Quantifying Cyber-Vulnerability in Power Electronics Systems via an Impedance-Based Attack Reachable Domain

ResearchDGX agent

arXiv:2605.14502v1 Announce Type: cross Abstract: Power electronics systems are increasingly exposed to cyber threats due to their integration with digital controllers and communication networks. Howe

Quantitative Video World Model Evaluation for Geometric-Consistency

SafetyDGX agent

arXiv:2605.15185v1 Announce Type: cross Abstract: Generative video models are increasingly studied as implicit world models, yet evaluating whether they produce physically plausible 3D structure and m

R2R2: Robust Representation for Intensive Experience Reuse via Redundancy Reduction in Self-Predictive Learning

SafetyDGX agent

arXiv:2605.14026v1 Announce Type: cross Abstract: For reinforcement learning in data-scarce domains like real-world robotics, intensive data reuse enhances efficiency but induces overfitting. While pr

REALM: Retrospective Encoder Alignment for LFP Modeling

Model ReleasesDGX agent

arXiv:2605.14867v1 Announce Type: cross Abstract: Spike activity has been the dominant neural signal for behavior decoding due to its high spatial and temporal resolution. However, as brain-computer i

ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing

ApplicationsDGX agent

arXiv:2507.21433v3 Announce Type: replace-cross Abstract: Large Reasoning Models (LRMs) are becoming integral to many AI inference systems, enhancing their capabilities with advanced reasoning. Howeve

Reinforcement Learning for Diffusion LLMs with Entropy-Guided Step Selection and Stepwise Advantages

SafetyDGX agent

arXiv:2603.12554v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has been effective for post-training autoregressive (AR) language models, but extending these methods to diffusion

Reinforcement Learning for Tool-Calling Agents in Fast Healthcare Interoperability Resources (FHIR)

Model ReleasesDGX agent

arXiv:2605.14126v1 Announce Type: cross Abstract: Fast Healthcare Interoperability Resources (FHIR) is the dominant standard for interoperable exchange of healthcare data. In FHIR, electronic health r

Residual Stream Duality in Modern Transformer Architectures

ResearchDGX agent

arXiv:2603.16039v2 Announce Type: replace-cross Abstract: Recent work has made clear that the residual pathway is not mere optimization plumbing; it is part of the model's representational machinery.

Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy

SafetyDGX agent

arXiv:2605.14558v1 Announce Type: cross Abstract: Agentic reinforcement learning trains large language models using multi-turn trajectories that interleave long reasoning traces with short environment

ROAD: Adaptive Data Mixing for Offline-to-Online Reinforcement Learning via Bi-Level Optimization

SafetyDGX agent

arXiv:2605.14497v1 Announce Type: cross Abstract: Offline-to-online reinforcement learning harnesses the stability of offline pretraining and the flexibility of online fine-tuning. A key challenge lie

Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons

Local AiDGX agent

arXiv:2603.02115v2 Announce Type: replace-cross Abstract: General-purpose robot reward models are typically trained to predict absolute task progress from expert demonstrations, providing only local,

ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety

Model ReleasesDGX agent

arXiv:2605.14152v1 Announce Type: cross Abstract: Safety evaluations for large language models (LLMs) increasingly target high-stakes National Security and Public Safety (NSPS) risks, yet multilingual

RQ-MoE: Residual Quantization via Mixture of Experts for Efficient Input-Dependent Vector Compression

ResearchDGX agent

arXiv:2605.14359v1 Announce Type: cross Abstract: Vector quantization is a fundamental tool for compressing high-dimensional embeddings, yet existing multi-codebook methods rely on static codebooks th

RxEval: A Prescription-Level Benchmark for Evaluating LLM Medication Recommendation

Model ReleasesDGX agent

arXiv:2605.14543v1 Announce Type: cross Abstract: Inpatient medication recommendation requires clinicians to repeatedly select specific medications, doses, and routes as a patient's condition evolves.

S-AI-Recursive: A Bio-Inspired and Temporal Sparse AI Architecture for Iterative, Introspective, and Energy-Frugal Reasoning

AgentsDGX agent

arXiv:2605.13872v1 Announce Type: cross Abstract: This article introduces S-AI-Recursive, a bio-inspired Sparse Artificial Intelligence architecture in which reasoning is operationalized as a hormonal

Safe Bayesian Optimization for Complex Control Systems via Additive Gaussian Processes

Model ReleasesDGX agent

arXiv:2408.16307v3 Announce Type: replace-cross Abstract: Automatic controller tuning is attractive for robotics and mechatronic systems whose dynamics are difficult to model accurately, but direct bl

Sat3DGen: Comprehensive Street-Level 3D Scene Generation from Single Satellite Image

Model ReleasesDGX agent

arXiv:2605.14984v1 Announce Type: cross Abstract: Generating a street-level 3D scene from a single satellite image is a crucial yet challenging task. Current methods present a stark trade-off: geometr

SceneFunRI: Reasoning the Invisible for Task-Driven Functional Object Localization

Model ReleasesDGX agent

arXiv:2605.14704v1 Announce Type: cross Abstract: In real-world scenes, target objects may reside in regions that are not visible. While humans can often infer the locations of occluded objects from c

Second-Order Actor-Critic Methods for Discounted MDPs via Policy Hessian Decomposition

SafetyDGX agent

arXiv:2605.14982v1 Announce Type: cross Abstract: We address the discounted reward setting in reinforcement learning (RL). To mitigate the value approximation challenges in policy gradient methods, ac

Seeing Through the Brain: New Insights from Decoding Visual Stimuli with fMRI

ApplicationsDGX agent

arXiv:2510.16196v2 Announce Type: replace-cross Abstract: Understanding how the brain encodes visual information is a central challenge in neuroscience and machine learning. A promising approach is to

Self-Distilled Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2605.15155v1 Announce Type: cross Abstract: Reinforcement learning (RL) has emerged as a central paradigm for post-training LLM agents, yet its trajectory-level reward signal provides only coars

Semantic Feature Segmentation for Interpretable Predictive Maintenance in Complex Systems

ResearchDGX agent

arXiv:2605.14318v1 Announce Type: new Abstract: Predictive maintenance in complex systems is often complicated by the heterogeneity and redundancy of monitored variables,which can obscure fault-releva

SemaTune: Semantic-Aware Online OS Tuning with Large Language Models

Model ReleasesDGX agent

arXiv:2605.15026v1 Announce Type: cross Abstract: Online OS tuning can improve long-running services, but existing controllers are poorly matched to live hosts. They treat scheduler, power, memory, an

Sequential Resource Trading Using Comparison-Based Gradient Estimation

AgentsDGX agent

arXiv:2408.11186v4 Announce Type: replace-cross Abstract: We study sequential multi-issue trading between two greedily rational agents who exchange resources from a finite set of categories. Each agen

Sheaf-Theoretic Transport and Obstruction for Detecting Scientific Theory Shift in AI Agents

Model ReleasesDGX agent

arXiv:2605.14033v1 Announce Type: new Abstract: Scientific theory shift in AI agents requires more than fitting equations to data. An artificial scientific agent must detect whether an existing repres

Silent Neuron Theory and Plasticity Preservation for Deep Reinforcement Learning in Adaptive Video Streaming

ApplicationsDGX agent

arXiv:2505.01584v3 Announce Type: replace-cross Abstract: Adaptive video streaming optimizes Quality of Experience (QoE) metrics by selecting appropriate bitrates according to varying network bandwidt

SimPersona: Learning Discrete Buyer Personas from Raw Clickstreams for Grounded E-Commerce Agents

SafetyDGX agent

arXiv:2605.14205v1 Announce Type: new Abstract: LLM-based web agents can navigate live storefronts, yet they often collapse to a single 'average buyer' policy, failing to capture the heterogeneous and

SkillFlow: Flow-Driven Recursive Skill Evolution for Agentic Orchestration

SafetyDGX agent

arXiv:2605.14089v1 Announce Type: new Abstract: In recent years, a variety of powerful LLM-based agentic systems have been applied to automate complex tasks through task orchestration. However, existi

SliceGraph: Mapping Process Isomers in Multi-Run Chain-of-Thought Reasoning

ResearchDGX agent

arXiv:2605.14619v1 Announce Type: new Abstract: Multi-run chain-of-thought reasoning is usually collapsed to final-answer aggregates, which discard howsampled trajectories share, split, and rejoin thr

Slot-MPC: Goal-Conditioned Model Predictive Control with Object-Centric Representations

SafetyDGX agent

arXiv:2605.14937v1 Announce Type: cross Abstract: Predictive world models enable agents to model scene dynamics and reason about the consequences of their actions. Inspired by human perception, object

Small, Private Language Models as Teammates for Educational Assessment Design

Local AiDGX agent

arXiv:2605.15015v1 Announce Type: new Abstract: Generative AI increasingly supports educational design tasks, e.g., through Large Language Models (LLMs), demonstrating the capability to design assessm

SparseOIT: Improving Order-Independent Transparency 3DGS via Active Set Method

ResearchDGX agent

arXiv:2605.13855v1 Announce Type: cross Abstract: 3D Gaussian Splatting (3DGS) has received tremendous popularity over the past few years due to its photorealistic visual appearance. However, 3DGS use

SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning

SafetyDGX agent

arXiv:2605.15044v1 Announce Type: cross Abstract: As audio-first agents become increasingly common in physical AI, conversational robots, and screenless wearables, audio large language models (audio-L

Spectral Analysis of Fake News Propagation

ApplicationsDGX agent

arXiv:2605.13861v1 Announce Type: cross Abstract: The propagation structure of fake news has been shown to be an important cue for detecting it; yet, existing propagation-based fake news detection met

SPIN: Structural LLM Planning via Iterative Navigation for Industrial Tasks

Model ReleasesDGX agent

arXiv:2605.14051v1 Announce Type: new Abstract: Industrial LLM agent systems often separate planning from execution, yet LLM planners frequently produce structurally invalid or unnecessarily long work

Spontaneous symmetry breaking and Goldstone modes for deep information propagation

ResearchDGX agent

arXiv:2605.14685v1 Announce Type: cross Abstract: In physical systems, whenever a continuous symmetry is spontaneously broken, the system possesses excitations called Goldstone modes, which allow cohe

Stateful Reasoning via Insight Replay

Model ReleasesDGX agent

arXiv:2605.14457v1 Announce Type: new Abstract: Chain-of-Thought (CoT) reasoning has become a foundation for eliciting multi-step reasoning in large language models, but recent studies show that its b

Streaming Speech-to-Text Translation with a SpeechLLM

ResearchDGX agent

arXiv:2605.14766v1 Announce Type: cross Abstract: Normally, a system that translates speech into text consists of separate modules for speech recognition and text-to-text translation. Combining those

SurgicalMamba: Dual-Path SSD with State Regramming for Online Surgical Phase Recognition

HardwareDGX agent

arXiv:2605.14889v1 Announce Type: cross Abstract: Online surgical phase recognition (SPR) underpins context-aware operating-room systems and requires committing to a prediction at every frame from pas

SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades

Model ReleasesDGX agent

arXiv:2605.14415v1 Announce Type: cross Abstract: Coding agents powered by large language models are increasingly expected to perform realistic software maintenance tasks beyond isolated issue resolut

Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks

Model ReleasesDGX agent

arXiv:2605.14604v1 Announce Type: new Abstract: This position paper argues that effective tutoring requires corrective friction: surfacing misconceptions and challenging them supportively to drive con

Synthesizing POMDP Policies: Sampling Meets Model-checking via Learning

SafetyDGX agent

arXiv:2605.14440v1 Announce Type: new Abstract: Partially Observable Markov Decision Processes (POMDPs) are the standard framework for decision-making under uncertainty. While sampling-based methods s

TAPIOCA: Why Task- Aware Pruning Improves OOD model Capability

ResearchDGX agent

arXiv:2605.14738v1 Announce Type: cross Abstract: Recent work has promoted task-aware layer pruning as a way to improve model performance on particular tasks, as shown by TALE. In this paper, we inves

TeachAnything: A Multimodal Crowdsourcing Platform for Training Embodied AI Agents in Symmetrical Reality

AgentsDGX agent

arXiv:2605.14556v1 Announce Type: new Abstract: Symmetrical Reality (SR) is emerging as a future trend for human-agent coexistence, placing higher demands on agents to acquire human-like intelligence.

Teaching and Evaluating LLMs to Reason About Polymer Design Related Tasks

Model ReleasesDGX agent

arXiv:2601.16312v2 Announce Type: replace-cross Abstract: Research in AI4Science has shown promise in many science applications, including polymer design. However, current LLMs are ineffective in this

Teaching Large Language Models When Not to Know: Learning Temporal Critique for Ex-Ante Reasoning

ResearchDGX agent

arXiv:2605.14636v1 Announce Type: new Abstract: Large language models (LLMs) often fail to reason under temporal cutoffs: when prompted to answer from the standpoint of an earlier time, they exploit k

TERMINATOR: Learning Optimal Exit Points for Early Stopping in Chain-of-Thought Reasoning

ResearchDGX agent

arXiv:2603.12529v2 Announce Type: replace-cross Abstract: Large Reasoning Models (LRMs) achieve impressive performance on complex reasoning tasks via Chain-of-Thought (CoT) reasoning, which enables th

TERMS-Bench: Diagnosing LLM Negotiation Agents Beyond Deal Rate

SafetyDGX agent

arXiv:2605.13909v1 Announce Type: cross Abstract: Negotiation is a central mechanism of economic exchange, shaping markets, procurement, labor agreements, and resource allocation. It is also a canonic

Text Knows What, Tables Know When: Clinical Timeline Reconstruction via Retrieval-Augmented Multimodal Alignment

Model ReleasesDGX agent

arXiv:2605.15168v1 Announce Type: cross Abstract: Reconstructing precise clinical timelines is essential for modeling patient trajectories and forecasting risk in complex, heterogeneous conditions lik

TFGN: Task-Free, Replay-Free Continual Pre-Training Without Catastrophic Forgetting at LLM Scale

Model ReleasesDGX agent

arXiv:2605.15053v1 Announce Type: cross Abstract: Continually pre-training a large language model on heterogeneous text domains, without replay or task labels, has remained an unsolved architectural p

The Evaluation Trap: Benchmark Design as Theoretical Commitment

Model ReleasesDGX agent

arXiv:2605.14167v1 Announce Type: new Abstract: Every AI benchmark operationalizes theoretical assumptions about the capability it claims to assess. When assumptions function as unexamined commitments

The Great Pretender: A Stochasticity Problem in LLM Jailbreak

SafetyDGX agent

arXiv:2605.14418v1 Announce Type: cross Abstract: 'Oh-Oh, yes, I'm the great pretender. Pretending that I'm doing well. My need is such, I pretend too much...' summarizes the state in the area of jail

The Impact of Generative AI on Collaborative Open-Source Software Development: Evidence from GitHub Copilot

ApplicationsDGX agent

arXiv:2410.02091v3 Announce Type: replace-cross Abstract: Generative artificial intelligence (AI) facilitates content production and enhances ideation capabilities, which can significantly influence d

The Moltbook Observatory Archive: an incremental dataset of agent-only social network activity

Model ReleasesDGX agent

arXiv:2605.13860v1 Announce Type: cross Abstract: Moltbook is a social media platform in which posts and comments are authored exclusively by autonomous AI agents. We present the Moltbook Observatory

← Previous
1…254255256257258…358
Next →