AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
11 May 2026

Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs

SafetyDGX agent

arXiv:2605.07447v1 Announce Type: cross Abstract: Vision-language models (VLMs) have advanced rapidly and are increasingly deployed in real-world applications, especially with the rise of agent-based

SparseRL-Sync: Lossless Weight Synchronization with ~100x Less Communication

SafetyDGX agent

arXiv:2605.07330v1 Announce Type: cross Abstract: In large-scale reinforcement learning (RL) systems with decoupled Trainer-Rollout execution, the Trainer must regularly synchronize policy weights to

Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer

ResearchDGX agent

arXiv:2605.07870v1 Announce Type: cross Abstract: We study the evolution of hidden-weight spectra in wide neural networks trained by (stochastic) gradient descent. We develop a two-level dynamical mea


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Spectral Filtering for Complex Linear Dynamical Systems

ResearchDGX agent

arXiv:2601.22400v2 Announce Type: replace-cross Abstract: We study the problem of learning complex-valued linear dynamical systems (CLDS) with sector-bounded spectrum. This class captures oscillatory

SpikingBrain: Spiking Brain-inspired Large Models

HardwareDGX agent

arXiv:2509.05276v4 Announce Type: replace-cross Abstract: Mainstream Transformer-based large language models face major efficiency bottlenecks: training computation scales quadratically with sequence

SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios

Model ReleasesDGX agent

arXiv:2605.07161v1 Announce Type: new Abstract: AI agents are increasingly used to diagnose and mitigate failures in production systems, known as agentic Site Reliability Engineering (SRE). Current SR

Stabilized neural Hamilton--Jacobi--Bellman solvers: Error analysis and applications in model-based reinforcement learning

SafetyDGX agent

arXiv:2605.07116v1 Announce Type: cross Abstract: Physics-informed neural solvers offer a promising route to model-based reinforcement learning in continuous time, where optimal feedback synthesis is

State Representation and Termination for Recursive Reasoning Systems

AgentsDGX agent

arXiv:2605.06690v1 Announce Type: new Abstract: Recursive reasoning systems alternate between acquiring new evidence and refining an accumulated understanding. Two design choices are typically left im

Statistical inference with belief functions: A survey

TutorialsDGX agent

arXiv:2605.07908v1 Announce Type: cross Abstract: Belief functions are a powerful and popular framework for the mathematical characterisation of uncertainty, in particular in situations in which lack

STDA-Net: Spectrogram-Based Domain Adaptation for cross-dataset Sleep Stage Classification

SafetyDGX agent

arXiv:2605.06736v1 Announce Type: cross Abstract: Accurate sleep stage classification across datasets remains challenging due to variability in EEG channel montages, sampling rates, recording environm

SteelDefectX: A Multi-Form Vision-Language Dataset and Benchmark for Steel Surface Defect Analysis

Model ReleasesDGX agent

arXiv:2603.21824v2 Announce Type: replace-cross Abstract: Steel surface defect analysis is critical for industrial quality control, yet existing benchmarks rely primarily on label-only annotations, li

Structural Rationale Distillation via Reasoning Space Compression

ResearchDGX agent

arXiv:2605.07139v1 Announce Type: cross Abstract: When distilling reasoning from large language models (LLMs) into smaller ones, teacher rationales for similar problems often vary wildly in structure

Structured Role-Aware Policy Optimization for Multimodal Reasoning

SafetyDGX agent

arXiv:2605.07274v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR), especially with Group Relative Policy Optimization (GRPO), has shown strong potential for improvi

Supervised sparse auto-encoders for interpretable and compositional representations

SafetyDGX agent

arXiv:2602.00924v2 Announce Type: replace Abstract: Sparse auto-encoders (SAEs) have re-emerged as a prominent method for mechanistic interpretability, yet they face two significant challenges: the no

Switchcraft: AI Model Router for Agentic Tool Calling

AgentsDGX agent

arXiv:2605.07112v1 Announce Type: new Abstract: Agentic AI systems that invoke external tools are powerful but costly, leading developers to default to large models and overspend inference budgets. Mo

Switching-time bioprocess control with pulse-width-modulated optogenetics

ApplicationsDGX agent

arXiv:2511.22893v2 Announce Type: replace-cross Abstract: Biotechnology can benefit from dynamic control to improve production efficiency. In this context, optogenetics enables modulation of gene expr

Sword: Style-Robust World Models as Simulators via Dynamic Latent Bootstrapping for VLA Policy Post-Training

Model ReleasesDGX agent

arXiv:2605.07288v1 Announce Type: cross Abstract: The integration of Vision-Language-Action (VLA) models with World Models has gained increasing attention. One representative approach treats learned W

Sycophantic AI makes human interaction feel more effortful and less satisfying over time

ApplicationsDGX agent

arXiv:2605.07912v1 Announce Type: cross Abstract: Millions of people now turn to artificial intelligence (AI) systems for personal advice, guidance, and support. Such systems can be sycophantic, frequ

Tacit Knowledge Extraction via Logic Augmented Generation and Active Inference

ApplicationsDGX agent

arXiv:2605.07639v1 Announce Type: new Abstract: Tacit knowledge plays a central role in human expertise, yet it remains difficult to capture, formalize, and reuse in machine-interpretable form. This c

TAP: Two-Stage Adaptive Personalization of Multi-Task and Multi-Modal Foundation Models in Federated Learning

Local AiDGX agent

arXiv:2509.26524v3 Announce Type: replace-cross Abstract: In federated learning (FL), local personalization of models has received significant attention, yet personalized fine-tuning of foundation mod

TAVIS: A Benchmark for Egocentric Active Vision and Anticipatory Gaze in Imitation Learning

Model ReleasesDGX agent

arXiv:2605.07943v1 Announce Type: cross Abstract: Active vision -- where a policy controls its own gaze during manipulation -- has emerged as a key capability for imitation learning, with multiple ind

TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue Agent

Model ReleasesDGX agent

arXiv:2601.18700v2 Announce Type: replace Abstract: Emotional Support Conversation requires not only affective expression but also grounded instrumental support to provide trustworthy guidance. Howeve

TeamBench: Evaluating Agent Coordination under Enforced Role Separation

Model ReleasesDGX agent

arXiv:2605.07073v1 Announce Type: new Abstract: Agent systems often decompose a task across multiple roles, but these roles are typically specified by prompts rather than enforced by access controls.

Temporal Smoothness Doubly Robust Learning for Debiased Knowledge Tracing

SafetyDGX agent

arXiv:2605.05958v2 Announce Type: replace Abstract: Knowledge Tracing (KT) is fundamental to intelligent education systems, yet relies on educational logs that are selectively observed. The non-random

Test-Time Compute Games

Model ReleasesDGX agent

arXiv:2601.21839v2 Announce Type: replace-cross Abstract: Test-time compute has emerged as a promising strategy to enhance the reasoning abilities of large language models (LLMs). However, this strate

Text-to-CAD Evaluation with CADTests

Model ReleasesDGX agent

arXiv:2605.07807v1 Announce Type: cross Abstract: Text-to-CAD has recently emerged as an important task with the potential to substantially accelerate design workflows. Despite its significance, there

The AI-Native Large-Scale Agile Software Development Manifesto

AgentsDGX agent

arXiv:2605.07717v1 Announce Type: cross Abstract: Despite the widespread adoption of agile methods, achieving true agility at scale remains elusive. Large-scale agile frameworks remain largely human-c

The Context Gathering Decision Process: A POMDP Framework for Agentic Search

AgentsDGX agent

arXiv:2605.07042v1 Announce Type: new Abstract: Large Language Model (LLM) agents are deployed in complex environments -- such as massive codebases, enterprise databases, and conversational histories

The EDelta-MHC-Geo Transformer: Adaptive Geodesic Operations with Guaranteed Orthogonality

Model ReleasesDGX agent

arXiv:2605.06729v1 Announce Type: cross Abstract: We present the EDelta-MHC-Geo Transformer, a novel architecture that unifies Manifold-Constrained Hyper-Connections (mHC), Deep Delta Learning (DDL),

The Effect of Mini-Batch Noise on the Implicit Bias of Adam

SafetyDGX agent

arXiv:2602.01642v2 Announce Type: replace-cross Abstract: With limited high-quality data and growing compute, multi-epoch training is gaining back its importance across sub-areas of deep learning. Ada

The Effective Depth Paradox: Evaluating the Relationship between Architectural Topology and Trainability in Deep CNNs

ResearchDGX agent

arXiv:2602.13298v3 Announce Type: replace-cross Abstract: This paper investigates the relationship between convolutional neural network (CNN) topology and image recognition performance through a compa

The Endogeneity of Miscalibration: Impossibility and Escape in Scored Reporting

SafetyDGX agent

arXiv:2605.07671v1 Announce Type: cross Abstract: Eliciting truthful reports from autonomous agents is a core problem in scalable AI oversight: a principal scores the agent's report using a strictly p

The Limits of AI-Driven Allocation: Optimal Screening under Aleatoric Uncertainty

SafetyDGX agent

arXiv:2605.07979v1 Announce Type: new Abstract: The rise of machine learning has shifted targeted resource allocation in policy and humanitarian settings toward algorithmic targeting based on predicte

The Memory Curse: How Expanded Recall Erodes Cooperative Intent in LLM Agents

Model ReleasesDGX agent

arXiv:2605.08060v1 Announce Type: cross Abstract: Context window expansion is often treated as a straightforward capability upgrade for LLMs, but we find it systematically fails in multi-agent social

The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment

SafetyDGX agent

arXiv:2605.07462v1 Announce Type: cross Abstract: Moltbook is a Reddit-like platform where OpenClaw agents post, comment, and vote at scale - a so far unprecedented incident that comes with serious sa

The Single-File Test: A Longitudinal Public-Interface Evaluation of First-Output LLM Web Generation with Social Reach Tracking

Model ReleasesDGX agent

arXiv:2605.06707v1 Announce Type: cross Abstract: This paper presents an eight-week observational comparison of 68 single-file HTML generations collected across 17 public experiments in the 'HTML AI B

The Text Uncanny Valley: Non-Monotonic Performance Degradation in LLM Information Retrieval

Model ReleasesDGX agent

arXiv:2605.07186v1 Announce Type: cross Abstract: Existing Large Language Model (LLM) benchmarks primarily focus on syntactically correct inputs, leaving a significant gap in evaluation on imperfect t

The Translation Tax Is Not a Scalar: A Counterfactual Audit of English-Source Cue Inheritance in Chinese Multilingual Benchmarks

Model ReleasesDGX agent

arXiv:2605.07093v1 Announce Type: cross Abstract: The Translation Tax is often treated as a scalar: translated benchmarks are assumed to inflate scores by preserving English-source cues. We audit this

Think Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous Vehicles

Local AiDGX agent

arXiv:2512.03454v4 Announce Type: replace-cross Abstract: Interpreting natural-language commands to localize target objects is critical for autonomous driving (AD). Existing visual grounding (VG) meth

THINKSAFE: Self-Generated Safety Alignment for Reasoning Models

Model ReleasesDGX agent

arXiv:2601.23143v2 Announce Type: replace Abstract: Large reasoning models (LRMs) achieve remarkable performance by leveraging reinforcement learning (RL) on reasoning tasks to generate long chain-of-

Three-in-One World Model: Energy-Based Consistency, Prediction, and Counterfactual Inference for Marketing Intervention

ResearchDGX agent

arXiv:2605.07199v1 Announce Type: new Abstract: Marketing decisions reflect the interaction of latent consumer heterogeneity, time-varying internal states, and explicit interventions, a structure that

TimeLesSeg: Unified Contrast-Agnostic Cross-Sectional and Longitudinal MS Lesion Segmentation via a Stochastic Generative Model

ResearchDGX agent

arXiv:2605.07955v1 Announce Type: cross Abstract: Multiple sclerosis (MS) expresses substantial clinical and radiological heterogeneity, which poses significant challenges for automatic lesion segment

Toeplitz MLP Mixers are Low Complexity, Information-Rich Sequence Models

Model ReleasesDGX agent

arXiv:2605.06683v1 Announce Type: cross Abstract: Transformer-based large language models are in some respects limited by the quadratic time and space computational complexity of attention. We introdu

Tool Calling is Linearly Readable and Steerable in Language Models

Model ReleasesDGX agent

arXiv:2605.07990v1 Announce Type: cross Abstract: When a tool-calling agent picks the wrong tool, the failure is invisible until execution: the email gets sent, the meeting gets missed. Probing 12 ins

Tools as Continuous Flow for Evolving Agentic Reasoning

Model ReleasesDGX agent

arXiv:2605.07339v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in orchestrating tools for reasoning tasks. However, existing methods rely on a s

TopoPrune: Robust Data Pruning via Unified Latent Space Topology

ResearchDGX agent

arXiv:2602.02739v2 Announce Type: replace-cross Abstract: Geometric data pruning methods, while practical for leveraging pretrained models, are fundamentally unstable. Their reliance on extrinsic geom

Toward Privileged Foundation Models:LUPI for Accelerated and Improved Learning

ResearchDGX agent

arXiv:2605.07799v1 Announce Type: cross Abstract: Training foundation models is computationally intensive and often slow to converge.We introduce PIQL,Privileged Information for Quick and Quality Lear

Towards an Inferentialist Account of Information Through Proof-theoretic Semantics

ResearchDGX agent

arXiv:2605.05368v2 Announce Type: replace-cross Abstract: Information is one of the most widely-discussed concepts of the current era. However, a great deal of insightful work notwithstanding, it is y

Towards Apples to Apples for AI Evaluations: From Real-World Use Cases to Evaluation Scenarios

ApplicationsDGX agent

arXiv:2605.07986v1 Announce Type: cross Abstract: AI measurement science has a wide variety of methodologies and measurements for comparing AI systems, resulting in what often appear to be 'apples-to-

Towards Autonomous Business Intelligence via Data-to-Insight Discovery Agent

AgentsDGX agent

arXiv:2605.07202v1 Announce Type: new Abstract: Transforming fragmented enterprise data into actionable insights remains a significant challenge for LLMs, constrained by complex database schemas, limi

Towards Billion-scale Multi-modal Biometric Search

HardwareDGX agent

arXiv:2605.07655v1 Announce Type: cross Abstract: Searching a multi-biometric database of a billion records for a country-level identity system requires pushing the limits of all aspects of a biometri

Towards Differentially Private Reinforcement Learning with General Function Approximation

SafetyDGX agent

arXiv:2605.07049v1 Announce Type: cross Abstract: We present the first theoretical guarantees for differentially private online reinforcement learning (RL) with general function approximation, extendi

Towards Security-Auditable LLM Agents: A Unified Graph Representation

AgentsDGX agent

arXiv:2605.06812v1 Announce Type: new Abstract: LLM-based agentic systems are rapidly evolving to perform complex autonomous tasks through dynamic tool invocation, stateful memory management, and mult

TRACE: Tourism Recommendation with Accountable Citation Evidence

ResearchDGX agent

arXiv:2605.07677v1 Announce Type: cross Abstract: Tourism is a high-stakes setting for conversational recommender systems (CRS): a plausible-sounding suggestion can waste real money and trip time once

TraceFix: Repairing Agent Coordination Protocols with TLA+ Counterexamples

AgentsDGX agent

arXiv:2605.07935v1 Announce Type: new Abstract: We present TraceFix, a verification-first pipeline for Large Language Model (LLM) multi-agent coordination. An agent synthesizes a protocol topology as

Tracing Uncertainty in Language Model 'Reasoning'

Model ReleasesDGX agent

arXiv:2605.07776v1 Announce Type: cross Abstract: Language model (LM) 'reasoning', commonly described as Chain-of-Thought or test-time scaling, often improves benchmark performance, but the dynamics u

Tracking Large-scale Shared Bikes with Inertial Motion Learning in GNSS Blocked Environments

Local AiDGX agent

arXiv:2605.07412v1 Announce Type: cross Abstract: Although Global Navigation Satellite Systems (GNSS) provide a general solution for bike tracking outdoors, there still exist complex riding environmen

Trajectory as the Teacher: Few-Step Discrete Flow Matching via Energy-Navigated Distillation

Model ReleasesDGX agent

arXiv:2605.07924v1 Announce Type: cross Abstract: Discrete flow matching generates text by iteratively transforming noise tokens into coherent language, but may require hundreds of forward passes. Dis

TSRBench: A Comprehensive Multi-task Multi-modal Time Series Reasoning Benchmark for Generalist Models

Model ReleasesDGX agent

arXiv:2601.18744v2 Announce Type: replace Abstract: Time series are ubiquitous in real-world scenarios and crucial for applications ranging from energy management to traffic control. Consequently, the

TTF: Temporal Token Fusion for Efficient Video-Language Model

ResearchDGX agent

arXiv:2605.07355v1 Announce Type: cross Abstract: Video-language models (VLMs) face rapid inference costs as visual token counts scale with video length. For example, 32 frames at 448{imes}448 resolut

← Previous
1…283284285286287…358
Next →