AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
12 May 2026

Privacy-Preserving Federated Learning: Integrating Zero-Knowledge Proofs in Scalable Distributed Architectures

Local AiDGX agent

arXiv:2605.08152v1 Announce Type: cross Abstract: The intersection of Artificial Intelligence (AI) and distributed systems has given rise to Federated Learning (FL), a paradigm that enables decentrali

ProactBench: Beyond What The User Asked For

Model ReleasesDGX agent

arXiv:2605.09228v1 Announce Type: cross Abstract: Most LLM benchmarks score how well a model responds to explicit requests. They leave unmeasured a different conversational ability: noticing and actin

Probing Cross-modal Information Hubs in Audio-Visual LLMs

ResearchDGX agent

arXiv:2605.10815v1 Announce Type: new Abstract: Audio-visual large language models (AVLLMs) have recently emerged as a powerful architecture capable of jointly reasoning over audio, visual, and textua


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Probing Routing-Conditional Calibration in Attention-Residual Transformers

ResearchDGX agent

arXiv:2605.09850v1 Announce Type: cross Abstract: Post-hoc calibration is usually evaluated as a function of logits or softmax confidence alone, even as routing-augmented architectures increasingly ac

Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark

Model ReleasesDGX agent

arXiv:2509.26574v4 Announce Type: replace Abstract: While large language models (LLMs) with reasoning capabilities are progressing rapidly on high-school math competitions and coding, can they reason

Probing the Impact of Scale on Data-Efficient, Generalist Transformer World Models for Atari

Model ReleasesDGX agent

arXiv:2605.08578v1 Announce Type: cross Abstract: Developing generalist systems that retain human-like data efficiency is a central challenge. While world models (WMs) offer a promising path, existing

Process Matters more than Output for Distinguishing Humans from Machines

Model ReleasesDGX agent

arXiv:2605.06524v2 Announce Type: replace Abstract: Reliable human-machine discrimination is becoming increasingly important as large language models and autonomous agents are deployed in online setti

Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions

Model ReleasesDGX agent

arXiv:2605.10664v1 Announce Type: cross Abstract: Activation steering controls language model behavior by adding directions to internal representations at inference time, but standard residual-stream

PromptDx: Differentiable Prompt Tuning for Multimodal In-Context Alzheimer's Diagnosis

ResearchDGX agent

arXiv:2605.08585v1 Announce Type: cross Abstract: Deep learning models in medical imaging typically operate as parametric memory, diagnosing patients by recalling fixed knowledge learned during traini

PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models

SafetyDGX agent

arXiv:2501.03544v5 Announce Type: replace-cross Abstract: Recent text-to-image (T2I) models have exhibited remarkable performance in generating high-quality images from text descriptions. However, the

Prospective Compression in Human Abstraction Learning

ApplicationsDGX agent

arXiv:2605.09985v1 Announce Type: new Abstract: A core challenge in program synthesis is online library learning: the incremental acquisition of reusable abstractions under uncertainty about future ta

ProteinOPD: Towards Effective and Efficient Preference Alignment for Protein Design

SafetyDGX agent

arXiv:2605.10189v1 Announce Type: cross Abstract: Designing proteins with desired functions or properties represents a core goal in synthetic biology and drug discovery. Recent advances in protein lan

Provable Anytime Ensemble Sampling Algorithms in Nonlinear Contextual Bandits

ResearchDGX agent

arXiv:2510.10730v2 Announce Type: replace-cross Abstract: We provide a unified algorithmic framework for ensemble sampling in nonlinear contextual bandits and develop corresponding regret bounds for t

Provable Sparse Inversion and Token Relabel Enhanced One-shot Federated Learning with ViTs

ResearchDGX agent

arXiv:2605.10748v1 Announce Type: cross Abstract: One-Shot Federated Learning, where a central server learns a global model in a single communication round, has emerged as a promising paradigm. Howeve

PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning

TutorialsDGX agent

arXiv:2605.09931v1 Announce Type: cross Abstract: Tool-integrated reasoning (TIR) enables large language models (LLMs) to enhance their capabilities by interacting with external tools, such as code in

Pseudo-Deliberation in Language Models: When Reasoning Fails to Align Values and Actions

SafetyDGX agent

arXiv:2605.09893v1 Announce Type: cross Abstract: Large language models (LLMs) are often evaluated based on their stated values, yet these do not reliably translate into their actions, a discrepancy t

PYTHALAB-MERA: Validation-Grounded Memory, Retrieval, and Acceptance Control for Frozen-LLM Coding Agents

Local AiDGX agent

arXiv:2605.08468v1 Announce Type: cross Abstract: Local LLM-based coding agents increasingly work in settings where correctness is earned through execution feedback, persistent state, and bounded repa

Q-learning with Adjoint Matching

SafetyDGX agent

arXiv:2601.14234v3 Announce Type: replace-cross Abstract: We propose Q-learning with Adjoint Matching (QAM), a novel TD-based reinforcement learning (RL) algorithm that tackles a long-standing challen

Quantile Geometry Regularization for Distributional Reinforcement Learning

TutorialsDGX agent

arXiv:2605.08182v1 Announce Type: cross Abstract: Quantile-based distributional reinforcement learning methods learn return distributions through sampled quantile regression, but their bootstrapped ta

Qwen Goes Brrr: Off-the-Shelf RAG for Ukrainian Multi-Domain Document Understanding

Model ReleasesDGX agent

arXiv:2605.10296v1 Announce Type: cross Abstract: We participated in the Fifth UNLP shared task on multi-domain document understanding, where systems must answer Ukrainian multiple-choice questions fr

RADAR: Redundancy-Aware Diffusion for Multi-Agent Communication Structure Generation

AgentsDGX agent

arXiv:2605.09907v1 Announce Type: new Abstract: Compared with individual agents, large language model based multi-agent systems have shown great capabilities consistently across diverse tasks, includi

RAG-HAR: Retrieval Augmented Generation-based Human Activity Recognition

ApplicationsDGX agent

arXiv:2512.08984v2 Announce Type: replace-cross Abstract: Human Activity Recognition (HAR) underpins applications in healthcare, rehabilitation, fitness tracking, and smart environments, yet existing

Randomized PCA Forest for Unsupervised Outlier Detection

ResearchDGX agent

arXiv:2508.12776v3 Announce Type: replace-cross Abstract: We propose a novel unsupervised outlier detection method based on Randomized Principal Component Analysis (PCA). Motivated by the performance

RAwR: Role-Aware Rewiring via Approximate Equitable Partition

Local AiDGX agent

arXiv:2605.09457v1 Announce Type: cross Abstract: While Graph Neural Networks (GNNs) have demonstrated significant efficacy in node classification tasks, where predictions rely on local neighborhood i

RDEx-CASK: Cauchy Mutation, Archive, and Stagnation Kick for RDEx-CSOP

Model ReleasesDGX agent

arXiv:2605.09652v1 Announce Type: cross Abstract: We extend RDEx-CSOP with 3 changes that target stagnation & late-stage variance, plus minor parameter tuning. The second scale factor in the standard

RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache

ResearchDGX agent

arXiv:2605.08317v1 Announce Type: cross Abstract: Large language models (LLMs) have shown strong performance across diverse tasks, but their inference with long input contexts is bottlenecked by memor

Re-Triggering Safeguards within LLMs for Jailbreak Detection

SafetyDGX agent

arXiv:2605.10611v1 Announce Type: cross Abstract: This paper proposes a jailbreaking prompt detection method for large language models (LLMs) to defend against jailbreak attacks. Although recent LLMs

Re^2Math: Benchmarking Theorem Retrieval in Research-Level Mathematics

Model ReleasesDGX agent

arXiv:2605.09012v1 Announce Type: new Abstract: Large language models are increasingly capable at closed-world mathematical reasoning, but research assistance also requires source-grounded use of the

Real vs. Semi-Simulated: Rethinking Evaluation for Treatment Effect Estimation

Model ReleasesDGX agent

arXiv:2605.10430v1 Announce Type: cross Abstract: Estimating heterogeneous treatment effects with machine learning has attracted substantial attention in both academic research and industrial practice

REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage

Model ReleasesDGX agent

arXiv:2604.01527v3 Announce Type: replace-cross Abstract: Production deployment of AI coding agents requires fast, reproducible evaluation signals. Existing industrial practices trade off speed and fi

REAP: Reinforcement-Learning End-to-End Autonomous Parking with Gaussian Splatting Simulator for Real2Sim2Real Transfer

AgentsDGX agent

arXiv:2605.08713v1 Announce Type: cross Abstract: In recent years, autonomous parking has made significant advances, yet parking tasks still face challenges in extreme scenarios such as mechanical and

Reasoning-Aware Training for Time Series Forecasting

ResearchDGX agent

arXiv:2605.08625v1 Announce Type: cross Abstract: Time Series Foundation Models (TSFMs) excel at numerical forecasting but operate as black boxes lacking qualitative reasoning. Conversely, applying LL

Reasoning Compression with Mixed-Policy Distillation

SafetyDGX agent

arXiv:2605.08776v1 Announce Type: new Abstract: Reasoning-centric large language models (LLMs) achieve strong performance by generating intermediate reasoning trajectories, but often incur excessive t

Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge

SafetyDGX agent

arXiv:2605.10805v1 Announce Type: new Abstract: Reasoning-capable large language models (LLMs) have recently been adopted as automated judges, but their benefits and costs in LLM-as-a-Judge settings r

Reconciling Consistency-Based Diagnosis with Actual-Causality-Based Explanations

ResearchDGX agent

arXiv:2605.08688v1 Announce Type: new Abstract: We establish, from the point of view of Explainable AI (XAI), connections between Consistency-Based Diagnosis (CBD), on one side, and Actual Causality a

Recovering Physical Dynamics from Discrete Observations via Intrinsic Differential Consistency

Model ReleasesDGX agent

arXiv:2605.08454v1 Announce Type: cross Abstract: Recovering continuous-time dynamics from discrete observations is difficult because local supervision (e.g., pointwise regression targets, derivative

Recursive Language Models

Model ReleasesDGX agent

arXiv:2512.24601v3 Announce Type: replace Abstract: We study allowing large language models (LLMs) to process arbitrarily long prompts through the lens of inference-time scaling. We propose Recursive

Regime-Calibrated Fleet Repositioning with a Spatial Queue-Regret Decomposition

AgentsDGX agent

arXiv:2604.03883v2 Announce Type: replace-cross Abstract: Ride-hailing and autonomous mobility-on-demand operators reposition idle supply before future demand is fully observed. We study a retrieval-c

REI-Bench: Can Embodied Agents Understand Vague Human Instructions in Task Planning?

Model ReleasesDGX agent

arXiv:2505.10872v4 Announce Type: replace-cross Abstract: Robot task planning decomposes human instructions into executable action sequences that enable robots to complete a series of complex tasks. A

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems

SafetyDGX agent

arXiv:2605.08378v1 Announce Type: cross Abstract: Reinforcement learning has become a powerful paradigm for improving the capability of intelligent systems, but its practical deployment faces two cent

Reinforcement Learning with Action Chunking

SafetyDGX agent

arXiv:2507.07969v4 Announce Type: replace-cross Abstract: We present Q-chunking, a simple yet effective recipe for improving reinforcement learning (RL) algorithms for long-horizon, sparse-reward task

Relational Retrieval: Leveraging Known-Novel Interactions for Generalized Category Discovery

SafetyDGX agent

arXiv:2605.09420v1 Announce Type: cross Abstract: In this study, we tackle Generalized Category Discovery (GCD) via a Relational Retrieval perspective, explicitly coupling labeled and unlabeled data t

Relations Are Channels: Knowledge Graph Embedding via Kraus Decompositions

SafetyDGX agent

arXiv:2605.10317v1 Announce Type: cross Abstract: Knowledge graph embedding (KGE) models typically represent each relation as an operator on entity embeddings. In this work, we identify three structur

Remember the Decision, Not the Description: A Rate-Distortion Framework for Agent Memory

AgentsDGX agent

arXiv:2605.10870v1 Announce Type: new Abstract: Long-horizon language agents must operate under limited runtime memory, yet existing memory mechanisms often organize experience around descriptive crit

Remix the Timbre: Diffusion-Based Style Transfer Across Polyphonic Stems

ResearchDGX agent

arXiv:2605.09259v1 Announce Type: cross Abstract: Timbre transfer aims to modify the timbral identity of a musical recording while preserving the original melody and rhythm. While single-instrument ti

ReplaySCM: A Benchmark for Executable Causal Mechanism Induction from Interventions

Model ReleasesDGX agent

arXiv:2605.08197v1 Announce Type: cross Abstract: Most causal benchmarks for language models score local answers or graph structure. We introduce ReplaySCM, a 1,300 item benchmark for executable causa

RePO-VLA: Recovery-Driven Policy Optimization for Vision-Language-Action Models

SafetyDGX agent

arXiv:2605.09410v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models remain brittle in long-horizon, contact-rich manipulation because success-only imitation provides little supervisi

Research on Security Enhancement Methods for Adversarial Robust Large Language Model Intelligent Agents for Medical Decision-Making Tasks

SafetyDGX agent

arXiv:2605.08257v1 Announce Type: cross Abstract: Motivated by the challenge to improve the adversarial robustness, security, and trust of medical decision making intelligent agents, this study develo

Resource-Aware Evolutionary Neural Architecture Search for Cardiac MRI Segmentation

ResearchDGX agent

arXiv:2605.08238v1 Announce Type: cross Abstract: Cardiac magnetic resonance (CMR) segmentation underpins quantitative assessment of ventricular structure and function, yet reliable delineation remain

Results and Retrospective Analysis of the CODS 2025 AssetOpsBench Challenge

AgentsDGX agent

arXiv:2605.08518v1 Announce Type: new Abstract: Competition retrospectives are useful when they explain what a leaderboard measured, how hidden evaluation changed conclusions, and which design pattern

Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient?

Model ReleasesDGX agent

arXiv:2605.10848v1 Announce Type: cross Abstract: Does a lexical retriever suffice as large language models (LLMs) become more capable in an agentic loop? This question naturally arises when building

Rethinking Constraint Awareness for Efficient State Embedding of Neural Routing Solver

ResearchDGX agent

arXiv:2605.10122v1 Announce Type: new Abstract: Heavy-Encoder-Light-Decoder (HELD) neural routing solvers have emerged as a promising paradigm due to their broad applicability across multiple vehicle

Rethinking Entropy Minimization in Test-Time Adaptation for Autoregressive Models

SafetyDGX agent

arXiv:2605.08186v1 Announce Type: cross Abstract: Test-Time Adaptation (TTA) via entropy minimization (EM) has proven effective for classification tasks, yet its application to generative autoregressi

Rethinking Evaluation of Multiple Sclerosis (MS) Lesion Segmentation Models

ApplicationsDGX agent

arXiv:2605.09666v1 Announce Type: cross Abstract: Multiple Sclerosis (MS) is a chronic autoimmune disease that can significantly reduce the quality of life of a patient. Existing treatment options can

Rethinking Gating Mechanism in Sparse MoE: Handling Arbitrary Modality Inputs with Confidence-Guided Gate

ApplicationsDGX agent

arXiv:2505.19525v3 Announce Type: replace-cross Abstract: Effectively managing missing modalities is a fundamental challenge in real-world multimodal learning scenarios, where data incompleteness ofte

Rethinking Loss Reweighting for Imbalance Learning as an Inverse Problem: A Neural Collapse Point of View

SafetyDGX agent

arXiv:2605.10047v1 Announce Type: cross Abstract: Loss reweighting is a widely used strategy for long-tailed classification, but existing reweighting strategies often rely on heuristics and rarely def

Rethinking Random Transformers as Adaptive Sequence Smoothers for Sleep Staging

Model ReleasesDGX agent

arXiv:2605.09905v1 Announce Type: cross Abstract: Automatic sleep staging commonly adopts Transformers under the assumption that they learn complex long-range dependencies. We challenge this view by r

Retrieve-then-Steer: Online Success Memory for Test-Time Adaptation of Generative VLAs

Model ReleasesDGX agent

arXiv:2605.10094v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models show strong potential for general-purpose robotic manipulation, yet their closed-loop reliability often degrades u

Revis: Sparse Latent Steering to Mitigate Object Hallucination in Large Vision-Language Models

ResearchDGX agent

arXiv:2602.11824v2 Announce Type: replace Abstract: Despite the advanced capabilities of Large Vision-Language Models (LVLMs), they frequently suffer from object hallucination. One reason is that visu

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic

ResearchDGX agent

arXiv:2605.09157v1 Announce Type: cross Abstract: Mixture policies theoretically offer greater flexibility than unimodal policies in continuous action reinforcement learning, but the practical benefit

← Previous
1…272273274275276…358
Next →