Kernel-Gradient Drifting Models
arXiv:2605.10727v1 Announce Type: new Abstract: We propose kernel-gradient drifting, a one-step generative modeling framework that replaces the fixed Euclidean displacement direction in drifting model
Knowledge catalogue
arXiv:2605.10727v1 Announce Type: new Abstract: We propose kernel-gradient drifting, a one-step generative modeling framework that replaces the fixed Euclidean displacement direction in drifting model
arXiv:2605.09213v1 Announce Type: cross Abstract: We study causal self-attention dynamics -- a toy model for decoder Transformers -- which we interpret as a non-exchangeable interacting particle syste
arXiv:2605.09487v1 Announce Type: new Abstract: Modern embodied agents achieve impressive performance, but their task knowledge is often stored in neural weights, latent state, or prompt-bound memory,
arXiv:2605.09299v1 Announce Type: cross Abstract: Reconstructing 3D fluid velocity fields from sparse 2D video observations is a highly ill-posed inverse problem, demanding both transport consistency
arXiv:2605.09994v1 Announce Type: cross Abstract: Modern Large Foundation Model (LFM) training has transformed the data pipeline from a static ingestion layer into a dynamic component that must co-evo
arXiv:2602.09297v2 Announce Type: replace Abstract: Transformers update token representations through multi-head attention and residual connections as X leftarrow X + sum_{i} P^{(i)}XW_{V_i}W_{o_i}, w
arXiv:2605.08755v1 Announce Type: new Abstract: Large reasoning models (LRMs) reach competition-level math and coding accuracy via long autoregressive decoding, making per-token decoding cost a primar
arXiv:2605.08626v1 Announce Type: cross Abstract: Large language models (LLMs) are transforming society, powering applications from smartphone assistants to autonomous driving. Yet cloud-based LLM ser
arXiv:2605.08732v1 Announce Type: cross Abstract: Modern vision-based world models can represent observations as compact yet expressive latent manifolds, but fast goal-oriented planning in these space
arXiv:2605.06366v2 Announce Type: replace Abstract: Diffusion language models (DLMs) have recently emerged as competitive alternatives to autoregressive (AR) language models, yet differences in their
arXiv:2605.09204v1 Announce Type: new Abstract: Backpropagation is inherently sequential across depth, creating an O(K)-deep dependency chain that bottlenecks parallel training. While parallel-scan fo
arXiv:2605.08552v1 Announce Type: cross Abstract: Independent Component Analysis (ICA) is a foundational tool for unsupervised representation learning, yet its high-dimensional theory remains largely
arXiv:2605.08209v1 Announce Type: new Abstract: Deep learning methods are widely used under diverse resource constraints, resulting in models of varying sizes, such as the Vision Transformer (ViT) ser
arXiv:2512.16875v4 Announce Type: replace-cross Abstract: We study the problem of finding confidence ellipsoids for an arbitrary distribution in high dimensions. Given samples from a distribution D an
arXiv:2605.09754v1 Announce Type: cross Abstract: Classical coding-theoretic guarantees often rely on trust assumptions, such as requiring sufficiently many honest nodes compared with adversarial ones
arXiv:2605.09993v1 Announce Type: new Abstract: Graph foundation models (GFMs), pretrained on massive graph data, have transformed graph machine learning by supporting general-purpose reasoning across
arXiv:2605.08506v1 Announce Type: new Abstract: Robust optimization (RO) provides a principled framework for decision-making under uncertainty, but its performance critically depends on the choice of
arXiv:2605.08958v1 Announce Type: new Abstract: Multiple technologies that measure expression levels of protein mixtures in the human body offer a potential for detection and understanding the disease
arXiv:2605.09019v1 Announce Type: cross Abstract: We extend quantum state tomography with minimal cumulative disturbance, first investigated in [arXiv:2406.18370], to arbitrary finite-dimensional pure
arXiv:2511.17994v2 Announce Type: replace Abstract: We study differentially private model training with stochastic gradient descent under learning rate scheduling and correlated noise. Although correl
arXiv:2605.09718v1 Announce Type: cross Abstract: Many systems in physics, engineering, and biology exhibit multiscale stochastic dynamics, where low-dimensional slow variables evolve under the influe
arXiv:2605.08211v1 Announce Type: cross Abstract: Channel-gain maps provide the channel gain between any two locations in a geographical region. They find numerous applications, from resource allocati
arXiv:2605.08811v1 Announce Type: cross Abstract: This paper investigates the learning theory of Transformer networks for regression tasks on the compact Euclidean domain [0,1]^d and d-dimensional com
arXiv:2605.09448v1 Announce Type: new Abstract: The transition to First-Price Auctions (FPA) in digital advertising has spurred significant research, yet existing work typically assumes access to a va
arXiv:2605.09818v1 Announce Type: new Abstract: Reinforcement learning (RL) in healthcare has had mixed results, with reward sparsity, unreliable off-policy evaluation, and deployment-simulation gap a
arXiv:2508.14137v2 Announce Type: replace Abstract: The Macroscopic Fundamental Diagram is a popular tool used to describe traffic dynamics in an aggregated way, with applications ranging from traffic
arXiv:2605.10151v1 Announce Type: new Abstract: This paper addresses the problem of learning to sparsify stochastic linear bandits, where a decision-maker sequentially selects actions from a high-dime
arXiv:2605.09183v1 Announce Type: new Abstract: Behavior cloning provides strong imitation learning guarantees when training and test environments share the same dynamics. However, in many deployment
arXiv:2601.21410v3 Announce Type: replace-cross Abstract: Large language models (LLMs) encode rich semantic knowledge that can be useful for supervised learning, but their outputs are unreliable as st
arXiv:2605.06394v1 Announce Type: cross Abstract: These lecture notes introduce some topics of classical statistical physics, particularly those that are relevant for neural networks and deep learning
arXiv:2509.25742v4 Announce Type: replace Abstract: Graph Contrastive Learning (GCL) has shown strong promise for unsupervised graph representation learning, yet its effectiveness on heterophilic grap
arXiv:2605.10810v1 Announce Type: new Abstract: We introduce an automatically generated benchmark for predicting hidden text in technical papers. A paper supplies visible context X and a hidden contin
arXiv:2509.20786v3 Announce Type: replace Abstract: Training deep neural networks with noise and data heterogeneity is a major challenge. We introduce Lightweight Learnable Adaptive Weighting (LiLAW),
arXiv:2505.17204v3 Announce Type: replace-cross Abstract: The sliced Wasserstein flow (SWF), a nonparametric and implicit generative gradient flow, is transformed into a Liouville partial differential
arXiv:2605.09518v1 Announce Type: new Abstract: Meta-learning for algorithm selection relies on a meta-dataset in which each row corresponds to a supervised learning dataset described by meta-features
arXiv:2605.10807v1 Announce Type: cross Abstract: The integration of Large Language Models (LLMs) into Electronic Design Automation (EDA) and hardware security is rapidly reshaping the semiconductor i
arXiv:2605.08850v1 Announce Type: cross Abstract: We design Local LMO - a new projection-free gradient-type method for constrained optimization. The key algorithmic idea is to replace the global linea
arXiv:2605.10777v1 Announce Type: new Abstract: The quality of open-weight language models has dramatically improved in recent years. Sharing weights greatly facilitates model adoption by enabling the
arXiv:2605.08996v1 Announce Type: new Abstract: Graph-based accelerators have been widely adopted in symbolic data processing applications such as genomics, cybersecurity, and artificial intelligence.
arXiv:2605.09649v1 Announce Type: new Abstract: The key-value (KV) cache is a major bottleneck in long-context inference, where memory and computation grow with sequence length. Existing KV eviction m
arXiv:2605.10196v1 Announce Type: new Abstract: High-throughput gene perturbation experiments can test several genetic interventions in parallel, yet experimental budgets remain limited. A central goa
arXiv:2605.10240v1 Announce Type: cross Abstract: Software vulnerability detection is critical for ensuring software security and reliability. Despite recent advances in deep learning, real-world vuln
arXiv:2605.10784v1 Announce Type: new Abstract: Multi-negative preference optimization under the Plackett--Luce (PL) model extends Direct Preference Optimization (DPO) by leveraging comparative signal
arXiv:2601.22320v2 Announce Type: replace Abstract: We study continual mean estimation, where data vectors arrive sequentially and the goal is to maintain accurate estimates of the running mean. We ad
arXiv:2605.08759v1 Announce Type: new Abstract: Existing granular-ball generation methods are still mainly driven by handcrafted quality measures and heuristic splitting or stopping criteria, which we
arXiv:2605.08777v1 Announce Type: cross Abstract: Mode separation, namely how sharply a distribution fragments into barrier-separated clusters, is a fundamental geometric property of densities, diffic
arXiv:2509.22196v2 Announce Type: replace Abstract: Disentangled representations seek to recover latent factors of variation underlying observed data, yet their identifiability is still not fully unde
arXiv:2602.02494v2 Announce Type: replace Abstract: Clinical brain-to-text interfaces are designed for paralysed patients who cannot provide extensive training recordings. Pre-training improves data-e
arXiv:2505.16741v4 Announce Type: replace Abstract: Minimum attention applies the least action principle to changes of control concerning state and time, first proposed by Brockett. The involved regul
arXiv:2605.08701v1 Announce Type: new Abstract: This data paper describes METBRA25Y, a harmonized archive of hourly surface meteorological observations from Brazil derived from public historical recor
arXiv:2605.09654v1 Announce Type: cross Abstract: Sampling from score-based diffusion models incurs bias due to both time discretisation and the approximation of the score function. A common strategy
arXiv:2605.08815v1 Announce Type: new Abstract: Predicting microbial operon co-membership requires integrating two complementary biological signals: protein-scale molecular identity and genome-context
arXiv:2605.09609v1 Announce Type: new Abstract: We provide a counterexample to the minimal unimodal conjecture for polynomial neural networks (PNNs) with power activation functions. Fixing the input a
arXiv:2605.10809v1 Announce Type: new Abstract: We investigate the learning task of language generation in the limit, but shift focus from the traditional time-of-last-mistake metric of a generator's
arXiv:2604.02438v2 Announce Type: replace Abstract: The deployment of reinforcement learning (RL)-based controllers on physical systems is often limited by poor generalization to real-world scenarios,
arXiv:2602.22611v2 Announce Type: replace Abstract: In Embedding-as-an-Interface (EaaI) settings, pre-trained models are queried for Intermediate Representations (IRs). The distributional properties o
arXiv:2502.20213v2 Announce Type: replace Abstract: Depression is a mental disorder and can cause a variety of symptoms, including psychological, physical, and social. Speech has been proved an object
arXiv:2605.08678v1 Announce Type: new Abstract: Modern AI progress has been driven by ML methods that are generalizable across settings and scalable to larger regimes. As large language models demonst
arXiv:2605.09724v1 Announce Type: new Abstract: Existing accounts of grokking explain the phenomena in terms of mechanistic frameworks such as circuit efficiency or lazy-to-rich transitions. However,
arXiv:2601.21266v3 Announce Type: replace Abstract: Neural network models are increasingly used for state estimation in control and decision-making, yet it remains unclear to what extent they behave a