AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
20 May 2026

The Insurability Frontier of AI Risk: Mapping Threats to Affirmative Coverage, Silent Exposures, and Exclusions

AgentsDGX agent

arXiv:2605.18784v1 Announce Type: cross Abstract: The rapid diffusion of agentic AI has created a new coverage problem for commercial insurance: some AI-mediated losses are now affirmatively insured,

The Routing and Filtering Structure of Attention

ResearchDGX agent

arXiv:2605.18826v1 Announce Type: cross Abstract: The attention interaction matrix QK^{op} contains two entangled computations: a skew-symmetric component that redistributes information between positi

The World Won't Stay Still: Programmable Evolution for Agent Benchmarks

Model ReleasesDGX agent

arXiv:2603.05910v2 Announce Type: replace Abstract: LLM-powered tool-calling agents fulfill user requests by interacting with environments, querying data, and invoking tools in a multi-turn process. Y


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Theory-optimal Quantization Based on Flatness

Model ReleasesDGX agent

arXiv:2605.18800v1 Announce Type: cross Abstract: Post-training quantization has emerged as a widely adopted technique for compressing and accelerating the inference of Large Language Models (LLMs). T

ThoughtTrace: Understanding User Thoughts in Real-World LLM Interactions

SafetyDGX agent

arXiv:2605.20087v1 Announce Type: cross Abstract: Conversational AI has now reached billions of users, yet existing datasets capture only what people say, not what they think. We introduce ThoughtTrac

To Call or Not to Call: Diagnosing Intrinsic Over-Calling Bias in LLM Agents

Model ReleasesDGX agent

arXiv:2605.18882v1 Announce Type: cross Abstract: LLM agents exhibit a consistent tendency to over-call, invoking tools even in situations where none is needed. On the When2Call benchmark, six models

Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models

ResearchDGX agent

arXiv:2605.19227v1 Announce Type: cross Abstract: Unified autoregressive models (UAMs) are transformer models that generate text as well as image tokens within a single autoregressive pass. Shared par

TORQ: Two-Level Orthogonal Rotation for MXFP4 Quantization

ResearchDGX agent

arXiv:2605.19561v1 Announce Type: cross Abstract: As Large Language Models (LLMs) advance toward practical deployment, the Microscaling FP4 (MXFP4) format has emerged as a cornerstone for next-generat

Toto 2.0: Time Series Forecasting Enters the Scaling Era

Model ReleasesDGX agent

arXiv:2605.20119v1 Announce Type: cross Abstract: We show that time series foundation models scale: a single training recipe produces reliable forecast-quality improvements from 4M to 2.5B parameters.

Toward an AI-Powered Computational Testbed for Workforce Policy

SafetyDGX agent

arXiv:2605.19064v1 Announce Type: cross Abstract: Workforce transformations are difficult to forecast and costly to mismanage. In particular, the integration of artificial intelligence into knowledge

Toward Training Superintelligent Software Agents through Self-Play SWE-RL

AgentsDGX agent

arXiv:2512.18552v2 Announce Type: replace-cross Abstract: While current software agents powered by large language models (LLMs) and agentic reinforcement learning (RL) can boost programmer productivit

Toward User Comprehension Supports for LLM Agent Skill Specifications

AgentsDGX agent

arXiv:2605.19362v1 Announce Type: cross Abstract: Users often interpret and select agent skills through their exttt{SKILL.md} specifications. To protect users, existing audits mainly focus on maliciou

Towards Family-Grouped Hierarchical Federated Learning on Sub-5KB Models: A Feasibility Study of Privacy-Preserving ECG Monitoring for Ultra-Resource-Constrained Wearables

ResearchDGX agent

arXiv:2605.18862v1 Announce Type: cross Abstract: Cardiovascular disease remains the leading cause of death worldwide, and early detection of arrhythmias through continuous ECG monitoring on wearable

Towards LLM-Assisted Architecture Recovery for Real-World ROS~2 Systems: An Agent-Based Multi-Level Approach to Hierarchical Structural Architecture Reconstruction

AgentsDGX agent

arXiv:2605.20055v1 Announce Type: cross Abstract: Explicit software architecture models are essential artifacts for communicating, analyzing, and evolving complex software-intensive systems. In ROS~2-

Towards Multi-Model LLM Schedulers: Empirical Insights into Offloading and Preemption

HardwareDGX agent

arXiv:2605.19593v1 Announce Type: new Abstract: Modern deployments of Large Language Models (LLMs) increasingly require serving multiple models with diverse architectures, sizes, and specialization on

Training Neural Networks with Optimal Double-Bayesian Learning

Model ReleasesDGX agent

arXiv:2605.20009v1 Announce Type: cross Abstract: Backpropagation with gradient descent is a common optimization strategy employed by most neural network architectures in machine learning. However, fi

Transformers Linearly Represent Highly Structured World Models

ResearchDGX agent

arXiv:2605.18847v1 Announce Type: cross Abstract: Do transformers, when trained on sequential reasoning traces, build internal models of the underlying task? And if so, does the structure of those int

Transforming Constraint Programs to Input for Local Search

ResearchDGX agent

arXiv:2605.19671v1 Announce Type: new Abstract: Applying local search algorithms to combinatorial optimization problems is not an easy feat. Typically, human intervention is required to compile the co

Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On

SafetyDGX agent

arXiv:2605.19035v1 Announce Type: new Abstract: The rapid advancement of Large Language Models has given rise to autonomous LLM-based agents capable of complex reasoning and execution. As these agents

TSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM Agents

SafetyDGX agent

arXiv:2602.11767v3 Announce Type: replace Abstract: Advances in large language models (LLMs) are driving a shift toward using reinforcement learning (RL) to train agents from iterative, multi-turn int

TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing

Model ReleasesDGX agent

arXiv:2605.18859v1 Announce Type: cross Abstract: LLM routing matters most in long-horizon applications such as coding agents, deep research systems, and computer-use agents, where a single user reque

Unlocking the Potential of Continual Model Merging: An ODE Perspective

Model ReleasesDGX agent

arXiv:2605.19409v1 Announce Type: cross Abstract: Continual Model Merging (CMM) enables rapid customization of foundation models across sequentially arriving tasks, offering a scalable alternative to

Using Aristotle API for AI-Assisted Theorem Proving in Lean 4: A Formalisation Case Study of the Grasshopper Problem

Local AiDGX agent

arXiv:2605.20120v1 Announce Type: new Abstract: AI-assisted theorem proving can now generate substantial Lean developments for olympiad-level mathematics, but the evidential status of such development

VCR: Learning Valid Contextual Representation for Incomplete Wearable Signals

ApplicationsDGX agent

arXiv:2605.18837v1 Announce Type: cross Abstract: Wearable devices enable continuous health monitoring from multimodal signals, but real-world deployment is hindered by limited labeled data and pervas

ViroGym: Realistic Large-Scale Benchmarks for Evaluating Viral Proteins

Model ReleasesDGX agent

arXiv:2603.06740v2 Announce Type: replace-cross Abstract: Protein language models (pLMs) have shown strong potential for zero-shot prediction of missense variant effects, yet systematic benchmarking o

VL-DPO: Vision-Language-Guided Finetuning for Preference-Aligned Autonomous Driving

AgentsDGX agent

arXiv:2605.20082v1 Announce Type: cross Abstract: The rapid growth of autonomous driving datasets has enabled the scaling of powerful motion forecasting models. While large-scale pretraining provides

WARC-Bench: Web Archive Based Benchmark for GUI Subtask Executions

Model ReleasesDGX agent

arXiv:2510.09872v2 Announce Type: replace-cross Abstract: Training web agents to navigate complex, real-world websites requires them to master extit{subtasks} - short-horizon interactions on multiple

What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents

AgentsDGX agent

arXiv:2605.19447v1 Announce Type: new Abstract: Reinforcement learning can train LLM agents from sparse task rewards, but long-horizon credit assignment remains challenging: a single success-or-failur

What Do Evolutionary Coding Agents Evolve?

Model ReleasesDGX agent

arXiv:2605.20086v1 Announce Type: cross Abstract: Recent work pairs LLMs with evolutionary search to iteratively generate, modify, and select code using task-specific feedback. These systems have prod

What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure Code

ResearchDGX agent

arXiv:2605.19762v1 Announce Type: new Abstract: Code has become a standard component of modern foundation language model (LM) training, yet its role beyond programming remains unclear. We revisit the

When Critics Disagree: Adaptive Reward Poisoning Attacks in RIS-Aided Wireless Control System

SafetyDGX agent

arXiv:2605.20037v1 Announce Type: cross Abstract: Reward-poisoning attacks present a significant risk to learning-based wireless control systems. Given this, we propose a Disagreement-Guided Reward Po

When Individually Calibrated Models Become Collectively Miscalibrated

AgentsDGX agent

arXiv:2605.18858v1 Announce Type: cross Abstract: Probabilistic prediction systems often aggregate probability estimates from multiple models into a single decision. A common assumption is that if eac

When Skills Don't Help: A Negative Result on Procedural Knowledge for Tool-Grounded Agents in Offensive Cybersecurity

AgentsDGX agent

arXiv:2605.20023v1 Announce Type: new Abstract: Agent Skills, structured packages of procedural knowledge loaded into an LLM agent at inference time, are widely reported to improve task pass rates by

When Tabular Foundation Models Meet Strategic Tabular Data: A Prior Alignment Approach

SafetyDGX agent

arXiv:2605.19662v1 Announce Type: new Abstract: Tabular foundation models based on pretrained prior-data fitted networks~(PFNs) have shown strong generalization on diverse tabular tasks, but they are

When the Majority Votes Wrong, the Intervention Timing for Test-Time Reinforcement Learning Hides in the Extinction Window

ResearchDGX agent

arXiv:2605.19444v1 Announce Type: cross Abstract: Test-time reinforcement learning (TTRL) reports substantial accuracy gains on mathematical reasoning benchmarks using majority vote as a pseudo-label

When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR

SafetyDGX agent

arXiv:2605.19425v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become the dominant paradigm for advanced reasoning in Large Language Models (LLMs), but rol

When Web Apps Heal Themselves: A MAPE-K Based Approach to Fault Tolerance and Adaptive Recovery

AgentsDGX agent

arXiv:2605.19261v1 Announce Type: cross Abstract: Ensuring the reliability and resilience of modern web applications remains a critical challenge due to increasing system complexity and dynamic runtim

Whispers of Wealth: Red-Teaming Google's Agent Payments Protocol via Prompt Injection

Model ReleasesDGX agent

arXiv:2601.22569v2 Announce Type: replace-cross Abstract: Large language model (LLM) based agents are increasingly used to automate financial transactions, yet their reliance on contextual reasoning e

WIND: Weather Inverse Diffusion for Zero-Shot Atmospheric Modeling

TutorialsDGX agent

arXiv:2602.03924v2 Announce Type: replace-cross Abstract: Deep learning has revolutionized weather forecasting, but many challenges remain, including climate modeling. Moreover, the current landscape

World-Ego Modeling for Long-Horizon Evolution in Hybrid Embodied Tasks

Model ReleasesDGX agent

arXiv:2605.19957v1 Announce Type: cross Abstract: World models are widely explored in embodied intelligence, yet they typically predict distinct evolutions of the world and the ego within a single str

YAC: Bridging Natural Language and Interactive Visual Exploration with Generative AI for Biomedical Data Discovery

AgentsDGX agent

arXiv:2509.19182v2 Announce Type: replace-cross Abstract: Incorporating natural language input has the potential to improve the capabilities of biomedical data discovery interfaces. However, user inte

ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language Models

Model ReleasesDGX agent

arXiv:2605.18879v1 Announce Type: cross Abstract: Large language models inevitably retain sensitive information, defined as inputs that may induce harmful generations, due to training on massive web c

19 May 2026

1GC-7RC: One Graphic Card -- Seven Research Challenges! How Good Are AI Agents at Doing Your Job?

Model ReleasesDGX agent

arXiv:2605.17046v1 Announce Type: cross Abstract: Autonomous AI coding agents are becoming a core tool for ML practitioners in industry and research alike. Despite this growing adoption, no standardiz

3DPhysVideo: Consistency-Guided Flow SDE for Video Generation via 3D Scene Reconstruction and Physical Simulation

Model ReleasesDGX agent

arXiv:2605.16795v1 Announce Type: cross Abstract: Video generative models have made remarkable progress, yet they often yield visual artifacts that violate grounding in physical dynamics. Recent works

A Conflict-aware Evidential Framework for Reliable Sleep Stage Classification

ApplicationsDGX agent

arXiv:2605.17021v1 Announce Type: new Abstract: Multi-view learning has been widely applied for sleep stage classification using multi-modal data. However, existing methods typically assume that diffe

A Distributional View for Visual Mechanistic Interpretability: KL-Minimal Soft-Constraint Principle

ResearchDGX agent

arXiv:2605.17504v1 Announce Type: cross Abstract: Most current paradigms in visual mechanistic interpretability (MI) remain confined to interpreting internal units of the vision model via heuristic me

A Global-Local Graph Attention Network for Traffic Forecasting

ApplicationsDGX agent

arXiv:2605.16726v1 Announce Type: new Abstract: Traffic forecasting is a significant part of intelligent transportation systems. One of the critical challenges of traffic forecasting is to find spatio

A Holistic Method for Superquadric Fitting Using Unsupervised Clustering Analysis

ResearchDGX agent

arXiv:2605.16779v1 Announce Type: cross Abstract: This work presents a novel method for fitting superquadrics to point clouds under the contamination of noise and outliers, which has many applications

A Machine Learning Framework for EEG-Based Prediction of Treatment Efficacy in Chronic Neck Pain

ApplicationsDGX agent

arXiv:2605.16326v1 Announce Type: cross Abstract: Chronic neck pain is a leading cause of disability worldwide, and current treatment selection remains largely trial and error. We present a machine le

A Machine With Human-Like Memory Systems

Model ReleasesDGX agent

arXiv:2204.01611v3 Announce Type: replace Abstract: Inspired by the cognitive science theory, we explicitly model an agent with both semantic and episodic memory systems, and show that it is better th

A Machine with Short-Term, Episodic, and Semantic Memory Systems

Model ReleasesDGX agent

arXiv:2212.02098v5 Announce Type: replace Abstract: Inspired by the cognitive science theory of the explicit human memory systems, we have modeled an agent with short-term, episodic, and semantic memo

A More Word-like Image Tokenization for MLLMs

ResearchDGX agent

arXiv:2605.17954v1 Announce Type: cross Abstract: Modern multimodal large language models (MLLMs) typically keep the language model fixed and train a visual projector that maps the pixels into a seque

A neurosymbolic Approach with Epistemic Deep Learning for Hierarchical Image Classification

ResearchDGX agent

arXiv:2605.16383v1 Announce Type: cross Abstract: Deep neural networks achieve high accuracy on image classification tasks. Yet, they often produce overconfident predictions as which fail to express e

A New Perspective on Precision and Recall for Generative Models

ResearchDGX agent

arXiv:2511.02414v2 Announce Type: replace Abstract: With the recent success of generative models in image and text, the question of their evaluation has recently gained a lot of attention. While most

A Practical Noise2Noise Denoising Pipeline for High-Throughput Raman Spectroscopy

ResearchDGX agent

arXiv:2605.18511v1 Announce Type: new Abstract: A lightweight and reproducible denoising pipeline for high-throughput Raman spectroscopy is presented. The approach relies on a one-dimensional convolut

A-ProS: Towards Reliable Autonomous Programming Through Multi-Model Feedback

Model ReleasesDGX agent

arXiv:2605.18073v1 Announce Type: cross Abstract: Large Language Models (LLMs) demonstrate strong potential for automated code generation, yet their ability to iteratively refine solutions using execu

A Scalable Tool for Measuring Manner and Result Verbs in Developmental Language Research

ResearchDGX agent

arXiv:2605.16654v1 Announce Type: cross Abstract: Manner and result verbs encode different aspects of event structure and have been discussed in developmental work as a potentially informative distinc

A Simplex Witness Certificate for Constant Collapse in Variational Autoencoders

SafetyDGX agent

arXiv:2605.18224v1 Announce Type: cross Abstract: This note studies exact constant collapse in variational autoencoders, where the encoder mean becomes independent of the input. The goal is to make th

A Structural Threshold in Decision Capacity Governs Collapse in Self-Play Reinforcement Learning

ResearchDGX agent

arXiv:2605.16315v1 Announce Type: cross Abstract: We show that a threshold in decision capacity determines whether self-play reinforcement learning agents collapse under asymmetric rule perturbations.

A Survey on Foundation Models for Personalized Federated Intelligence

Model ReleasesDGX agent

arXiv:2505.06907v2 Announce Type: replace Abstract: The rise of large language models (LLMs), such as ChatGPT, Gemini, and Grok, has reshaped the AI landscape. As prominent instances of foundational m

← Previous
1…231232233234235…358
Next →