Multi-Token Residual Prediction
arXiv:2605.18817v1 Announce Type: new Abstract: Diffusion Language Models (DLMs) generate text by iteratively denoising masked token sequences, offering a tradeoff between parallelism and quality comp
Knowledge catalogue
arXiv:2605.18817v1 Announce Type: new Abstract: Diffusion Language Models (DLMs) generate text by iteratively denoising masked token sequences, offering a tradeoff between parallelism and quality comp
arXiv:2404.16676v2 Announce Type: replace-cross Abstract: We establish Multilayer Correlation Clustering, a novel generalization of Correlation Clustering to the multilayer setting. In this model, we
arXiv:2507.06428v2 Announce Type: replace-cross Abstract: We mathematically analyze and numerically study an actor-critic machine learning algorithm for solving high-dimensional Hamilton-Jacobi-Bellma
arXiv:2503.04929v3 Announce Type: replace-cross Abstract: Planning and control for high-dimensional robot manipulators in cluttered dynamic environments require computational efficiency and robust saf
arXiv:2603.24400v2 Announce Type: replace-cross Abstract: We propose a neural network model for contextual regression in which the regression model depends on contextual features that determine the ac
arXiv:2605.17326v1 Announce Type: cross Abstract: We investigate the role of the noise schedule in diffusion processes on Lie groups, with particular emphasis on applications to lattice gauge theory.
arXiv:2605.19965v1 Announce Type: new Abstract: Blind source separation (BSS) is a natural framework for studying how latent causes may be recovered from sensory mixtures, but deriving online and biol
arXiv:2605.18825v1 Announce Type: new Abstract: Prefix caching is a key optimization in Large Language Model (LLM) serving, reusing attention Key-Value (KV) states across requests with shared prompt p
arXiv:2601.12238v4 Announce Type: replace-cross Abstract: In this paper, we provide a comprehensive theoretical analysis of Stochastic Gradient Descent (SGD) and its momentum variants (Polyak Heavy-Ba
arXiv:2605.19584v1 Announce Type: new Abstract: We study an online market-making problem in which a learner sequentially posts bid and ask prices for a single asset while interacting with traders hold
arXiv:2605.19625v1 Announce Type: new Abstract: We study the problem of reconstructing an unknown point in R^d from approximate linear queries. This setting arises naturally in applications ranging fr
arXiv:2605.20105v1 Announce Type: new Abstract: Learning to generalise from limited data is a fundamental challenge for both artificial and biological systems. A common strategy is to extract reusable
arXiv:2605.20122v1 Announce Type: cross Abstract: Squared Wasserstein distance is a frequently used tool to measure discrepancy between probability distributions. This distance is typically computed b
arXiv:2605.19107v1 Announce Type: new Abstract: Green hydrogen plays an essential role in decarbonization, with capacity projected to scale to 560 GW by 2030 (vs. 1.39 GW in 2023) in net-zero settings
arXiv:2605.19589v1 Announce Type: new Abstract: Dental aerosol procedures produce sub-50 micrometre nuclei that can remain airborne for long periods in enclosed clinics, creating pathways for airborne
arXiv:2605.19145v1 Announce Type: new Abstract: In the literature, many continual learning (CL) algorithms have been proposed to address the issue of catastrophic forgetting in ML models (i.e., learni
arXiv:2605.18893v1 Announce Type: new Abstract: Graph Neural Networks (GNNs) are powerful tools for learning from graph-structured data, but their scalability is increasingly strained by the size of r
arXiv:2605.19610v1 Announce Type: cross Abstract: We investigate the asymptotic properties of the Levy Adaptive B-spline (LABS) regression model, a Bayesian nonparametric method that incorporates B-sp
arXiv:2605.19208v1 Announce Type: cross Abstract: Physical activity (PA) plays an important role in maintaining and improving health. Daily steps have been a key PA measure that is easily accessible w
arXiv:2605.19685v1 Announce Type: cross Abstract: Accurately assessing financial risk requires capturing both individual asset volatility and the complex, asymmetric dependence structures that emerge
arXiv:2411.10959v4 Announce Type: replace-cross Abstract: We study causal inference in experiments and quasi-experiments, where the economic outcome is imperfectly measured by a remotely sensed variab
arXiv:2605.19549v1 Announce Type: cross Abstract: Deep neural networks (DNNs) are suffering from ethical issues such as individual discrimination. In response, extensive NN repair techniques have been
arXiv:2605.19052v1 Announce Type: cross Abstract: Lagrangian Relaxation (LR) is a powerful technique for solving large-scale Mixed Integer Linear Programming (MILP), particularly those with decomposab
arXiv:2605.18821v1 Announce Type: new Abstract: Machine learning has revolutionized numerous industrial domains. Despite recent advances, machine learning models remain vulnerable to adversarial threa
arXiv:2504.17548v2 Announce Type: replace-cross Abstract: Anomaly Detection (AD) defines the task of identifying observations or events that deviate from typical - or normal - patterns, a critical cap
arXiv:2503.22823v3 Announce Type: replace-cross Abstract: In classical information theory, the Doeblin coefficient of a classical channel provides an efficiently computable upper bound on the total-va
arXiv:2605.19233v1 Announce Type: cross Abstract: Unmanned aerial vehicles (UAVs) are cyber-physical systems whose attack surface spans networked avionics and on-board sensor fusion: a compromised GPS
arXiv:2605.19170v1 Announce Type: cross Abstract: Diffusion/score-based models have emerged as powerful generative models, capable of generating high-quality samples that mimic the training data distr
arXiv:2605.19282v1 Announce Type: new Abstract: Muon is a matrix-aware optimizer that leverages Newton-Schulz (NS) iterations to enforce spectral gradient orthogonalization by driving all singular val
arXiv:2412.02818v4 Announce Type: replace-cross Abstract: Robot manipulation policies, while central to the promise of physical AI, are highly vulnerable in the presence of external variations in the
arXiv:2510.09174v3 Announce Type: replace Abstract: This paper takes a closer look at Git Re-Basin, an interesting new approach to merge trained models. We propose a hierarchical model merging scheme
arXiv:2605.18842v1 Announce Type: new Abstract: Safe reinforcement learning in nonstationary environments requires safety mechanisms that adapt as environmental conditions change. Standard safe reinfo
arXiv:2605.19014v1 Announce Type: new Abstract: Microsimulation models used by ministries of finance and central banks rely on parametric processes for lifetime earnings that capture only first and se
arXiv:2605.20157v1 Announce Type: new Abstract: Music streaming fraud, where bad actors artificially inflate stream counts to manipulate chart rankings and royalty payments, poses a significant threat
arXiv:2510.18821v3 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) has become the mainstream technique for training LLM agents. However, RLVR highly depends on w
arXiv:2605.19193v1 Announce Type: new Abstract: Multi-agent LLM debate improves factuality and reasoning, but most recipes pick a fixed round count, over-spending on easy items and under-spending on h
arXiv:2605.19830v1 Announce Type: new Abstract: Conventional treatment policies map patient covariates to a single recommended intervention in order to maximize expected clinical outcomes. Although a
arXiv:2408.12385v3 Announce Type: replace-cross Abstract: We study the problem of approximately recovering a probability distribution given noisy measurements of its Chebyshev polynomial moments. This
arXiv:2605.20069v1 Announce Type: new Abstract: Competitive selection processes, from scientific funding to admissions and hiring, use evaluations to score candidates, and eventually choose a subset o
arXiv:2605.18835v1 Announce Type: new Abstract: Traditional sheet metal forming relies on time-consuming and expensive Finite Element Analysis (FEA) for design validation, a process that significantly
arXiv:2602.18718v2 Announce Type: replace-cross Abstract: For approximating a target distribution given only its unnormalized log-density, stochastic gradient-based variational inference (VI) algorith
arXiv:2605.18851v1 Announce Type: new Abstract: Recent advances in Reinforcement Learning (RL) have underscored its potential for incentivizing reasoning capabilities of Large Language Models (LLMs).
arXiv:2605.18979v1 Announce Type: new Abstract: We propose Tabular Q-Learning (TabQL), a reinforcement learning framework that replaces the conventional parametric Q-network in Deep Q-Learning (DQN) w
arXiv:2602.11910v2 Announce Type: replace-cross Abstract: Audio diffusion models can synthesize high-fidelity music from text, yet achieving fine-grained control over specific musical attributes remai
arXiv:2605.20068v1 Announce Type: cross Abstract: Standard generative models struggle with heavy-tailed data: Lipschitz architectures cannot produce power-law tails from Gaussian noise, and interpolat
arXiv:2605.20030v1 Announce Type: new Abstract: While optimal transport (OT) enforces a rigid constraint by requiring two measures to be matched exactly, partial optimal transport relaxes this require
arXiv:2603.07018v2 Announce Type: replace-cross Abstract: Treatment effects estimated from a randomized controlled trial are local not only to the study population but also to the time at which the tr
arXiv:2605.18843v1 Announce Type: new Abstract: Backtesting large language models on historical events requires reasoning exclusively from information available before a specified cutoff date. Yet mod
arXiv:2605.19076v1 Announce Type: new Abstract: Inferring unknown initial states in shock-dominated compressible flows from sparse and noisy measurements is a challenging ill-posed inverse problem due
arXiv:2605.19537v1 Announce Type: new Abstract: Progress in LLMs is increasingly measured through standardized benchmarks, where state-of-the-art improvements are often separated by fractions of a per
arXiv:2510.20035v3 Announce Type: replace-cross Abstract: Vine copulas offer flexible multivariate dependence modeling and have become widely used in machine learning. Yet, structure learning remains
arXiv:2605.19403v1 Announce Type: new Abstract: Recent Continuous Thought Machine architecture decouples internal computation from external inputs via neural dynamics, but relies on multi-layer percep
arXiv:2504.04349v3 Announce Type: replace-cross Abstract: We examine fixed-price mechanisms in bilateral trade through the lens of regret minimization. Our main results are twofold. (i) For independen
arXiv:2503.17581v2 Announce Type: replace-cross Abstract: A computational method for the synthesis of time-optimal feedback control laws for linear nilpotent systems is proposed. The method is based o
arXiv:2605.18831v1 Announce Type: cross Abstract: Cold-chain storage limits access to insulin for hundreds of millions of people; a thermally protective patch polymer could help, but the design space
arXiv:2605.20074v1 Announce Type: new Abstract: Distillation transfers knowledge from a large model trained on broad data to a smaller, more efficient model suitable for deployment. In structured pred
arXiv:2605.20028v1 Announce Type: new Abstract: Bayesian filtering is a well-known problem that aims to estimate plausible states of a dynamical system from observations. Among existing approaches to
arXiv:2605.20134v1 Announce Type: new Abstract: Learning generalizable trajectory representations from raw GPS traces remains difficult because the data is continuous, noisy, and irregularly sampled.
arXiv:2605.19391v1 Announce Type: cross Abstract: Diffusion models have achieved remarkable success in generating samples from unknown data distributions. Most popular stochastic differential equation
arXiv:2602.03839v2 Announce Type: replace Abstract: Bandwidth-constrained distributed reinforcement learning (RL) post-training of large language models is bottlenecked by two channels: weight synchro