Exact Sequence Interpolation with Transformers
arXiv:2502.02270v3 Announce Type: replace Abstract: We prove that transformers can exactly interpolate datasets of finite input sequences in R^d, dgeq 2, with corresponding output sequences of smaller
Knowledge catalogue
arXiv:2502.02270v3 Announce Type: replace Abstract: We prove that transformers can exactly interpolate datasets of finite input sequences in R^d, dgeq 2, with corresponding output sequences of smaller
arXiv:2505.18604v3 Announce Type: replace Abstract: State-Space Models (SSMs) excel at capturing long-range dependencies with structured recurrence, making them well-suited for sequence modeling. Howe
arXiv:2601.19979v2 Announce Type: replace-cross Abstract: We develop a reinforcement learning algorithm to study the holographic entropy cone. Given a target entropy vector, our algorithm searches for
arXiv:2605.12995v1 Announce Type: new Abstract: Traditional retrieval pipelines optimize utility through stages of candidate retrieval and reranking, where ranking operates over a predefined candidate
arXiv:2605.13759v1 Announce Type: new Abstract: Clustering is an unsupervised machine learning task that consists of identifying groups of similar objects. It has numerous applications and is increasi
arXiv:2409.02708v2 Announce Type: replace Abstract: Data scarcity poses a serious threat to modern machine learning and artificial intelligence, as their practical success typically relies on the avai
arXiv:2605.13170v1 Announce Type: new Abstract: Multi-agent systems rely on communication for information sharing and action coordination, which exposes a vulnerability to attacks. We investigate sing
arXiv:2602.06138v2 Announce Type: replace Abstract: Generative policies based on diffusion models and flow matching have shown strong promise for offline reinforcement learning (RL), but their applica
arXiv:2605.13788v1 Announce Type: new Abstract: Active learning for machine-learning interatomic potentials (MLIPs) must address several challenges to be practical: scaling to large candidate pools, l
arXiv:2605.12997v1 Announce Type: new Abstract: Neural operators learn to map initial conditions to the terminal solution of partial differential equations (PDEs), providing a surrogate for the full o
arXiv:2605.12788v1 Announce Type: new Abstract: Sustained effort is essential for realizing the benefits of intelligent tutoring systems (ITS), yet many learners disengage or underuse available practi
arXiv:2605.12944v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) data selection is commonly formulated as instance ranking: score each example and retain a top-k subset. However, effective
arXiv:2605.13217v1 Announce Type: cross Abstract: Reinforcement learning has become a powerful paradigm for post-training large language model agents, yet credit assignment in multi-turn environments
arXiv:2406.13619v4 Announce Type: replace-cross Abstract: This paper develops a generative model by minimizing the second-order Wasserstein loss (the W_2 loss) through a distribution-dependent ordinar
arXiv:2605.13150v1 Announce Type: cross Abstract: Discrete automated processes in industrial and cyber-physical systems often exhibit a repetitive structure in which successive repetitions follow a co
arXiv:2605.13352v1 Announce Type: new Abstract: Standard dual-encoder vision-language models that map images and text to deterministic points on a shared unit hypersphere through ell_2 normalization t
arXiv:2509.19929v4 Announce Type: replace-cross Abstract: Uncertainty Quantification (UQ) is paramount for inference in engineering. A common inference task is to recover full-field information of phy
arXiv:2601.11942v3 Announce Type: replace Abstract: Variational quantum circuits are increasingly studied as continuous-function approximators, but quantum regression remains difficult to train when g
arXiv:2605.13743v1 Announce Type: new Abstract: Open datasets and benchmarks for entity-level carbon-emission prediction remain fragmented across access, scale, granularity, and evaluation. We introdu
arXiv:2501.17443v4 Announce Type: replace Abstract: Existing machine learning literature lacks graph-based domain adaptation techniques capable of handling large distribution shifts, primarily due to
arXiv:2605.12782v1 Announce Type: new Abstract: Financial transaction fraud prevention faces challenges such as complex relationship structures, concealed behavioral patterns, and dynamically changing
arXiv:2605.13673v1 Announce Type: new Abstract: The multicut problem is an NP-hard combinatorial optimization problem with diverse applications in fields such as bioinformatics, data mining and comput
arXiv:2605.12823v1 Announce Type: new Abstract: Coarse-grained (CG) molecular dynamics enables simulations of atomic systems such as biomolecules at timescales inaccessible to all-atom (AA) methods, b
arXiv:2605.13343v1 Announce Type: cross Abstract: Neural preconditioners for real-time physics simulation offer promising data-driven priors, but they often fail to capture long-range couplings effici
arXiv:2505.14587v2 Announce Type: replace-cross Abstract: Bootstrap methods have long been the cornerstone of ensemble learning in machine learning. This paper presents a theoretical analysis of boots
arXiv:2601.01860v2 Announce Type: replace Abstract: Detecting high-order epistasis is a fundamental challenge in genetic association studies due to the combinatorial explosion of candidate locus combi
arXiv:2601.19208v2 Announce Type: replace-cross Abstract: Semantic associations such as the link between 'bird' and 'flew' are foundational for language modeling as they enable models to go beyond mem
arXiv:2602.11618v4 Announce Type: replace Abstract: Chemical Language Models (CLMs) pre-trained on large scale molecular data are widely used for molecular property prediction. However, the common bel
arXiv:2605.12785v1 Announce Type: new Abstract: Hybrid machine learning combines physical knowledge with data-driven models to enhance interpretability and performance. In this context, Port-Hamiltoni
arXiv:2605.12693v1 Announce Type: new Abstract: Decision-focused learning trains predictive models end-to-end against downstream decision loss, but online settings suffer delayed feedback: outcomes ma
arXiv:2605.12999v1 Announce Type: cross Abstract: Closed-loop brain-computer interfaces often require both a forecast of upcoming neural population activity and a readout of the animal's behavioral st
arXiv:2602.04923v2 Announce Type: replace Abstract: Neural operators have emerged as powerful surrogates for the solution of partial differential equations (PDEs), yet their ability to handle general,
arXiv:2605.12765v1 Announce Type: new Abstract: Large Language Models memorize vast amounts of training data, raising concerns regarding privacy, copyright infringement, and safety. Machine unlearning
arXiv:2605.13786v1 Announce Type: new Abstract: Background: Pregnancy-associated thrombotic microangiopathy (P-TMA) is rare but life-threatening. Early risk prediction before overt clinical presentati
arXiv:2605.12768v1 Announce Type: cross Abstract: Open time-series forecasting (TSF) benchmarks cover retail, energy, weather, and traffic, but supply-chain logistics remains underserved. We introduce
arXiv:2605.12924v1 Announce Type: new Abstract: The instrumental-variables (IV) setting is standard for partial identification of causal effects when unobserved confounding makes point identification
arXiv:2605.13013v1 Announce Type: new Abstract: Diffusion world models have recently become competitive for online model-based reinforcement learning, but current approaches expose a tension: pixel di
arXiv:2605.13133v1 Announce Type: new Abstract: While EEG foundation models have shown significant potential in universal neural decoding across tasks, their advancement remains constrained by the ina
arXiv:2605.13160v1 Announce Type: cross Abstract: Modern Bayesian optimization and adaptive sampling methods increasingly rely on nonlinear parametric models, yet theoretical guarantees for such model
arXiv:2505.04613v4 Announce Type: replace-cross Abstract: We prove that kernel covariance embeddings lead to information-theoretically perfect separation of distinct continuous probability distributio
arXiv:2605.13045v1 Announce Type: new Abstract: The existing methods for evaluating the medical knowledge of Large Language Models (LLMs) are largely based on atemporal examination-style benchmarks, w
arXiv:2605.12714v1 Announce Type: new Abstract: Hidden states change substantially across the layers of modern language models, but most layer-wise analyses focus on one aspect of that change. We prop
arXiv:2506.11274v2 Announce Type: replace-cross Abstract: Test-time scaling has emerged as an effective approach for improving language model performance by utilizing additional compute at inference t
arXiv:2605.13284v1 Announce Type: cross Abstract: Recent advancements in large language models demonstrate that injecting perturbations can substantially enhance extrapolation performance. However, cu
arXiv:2605.13740v1 Announce Type: new Abstract: Whether navigating a building, operating a robot, or playing a game, an agent that acts effectively in an environment must first learn an internal model
arXiv:2602.13155v2 Announce Type: replace Abstract: Neural networks, particularly message-passing neural networks (MPNNs), are increasingly used as heuristics for hard combinatorial optimization probl
arXiv:2605.12561v1 Announce Type: new Abstract: Safe reinforcement learning (RL) typically asks extit{what} an agent should do. We ask extit{when} it needs to act, and show that a single policy can jo
arXiv:2605.12741v1 Announce Type: new Abstract: Enabling Large Language Models (LLMs) to continuously improve from environmental interactions is a central challenge in post-training. While on-policy s
arXiv:2605.13424v1 Announce Type: new Abstract: We propose last-mile fine-tuning, or Lift, a pipeline in which a pre-trained large language model extracts an initial table from unstructured clipboard
arXiv:2605.13265v1 Announce Type: new Abstract: Split learning (SL) enables collaborative training by partitioning a neural network across clients and a central server, but the cut-layer interface int
arXiv:2511.10709v2 Announce Type: replace-cross Abstract: Machine learning models are used for pattern recognition analysis of big data, without direct human intervention. The task of unsupervised lea
arXiv:2605.13503v1 Announce Type: cross Abstract: A key technical difficulty in differential privacy is selecting a privacy budget that satisfies privacy requirements while maximizing utility. A natur
arXiv:2601.06147v2 Announce Type: replace Abstract: Recent work has demonstrated surprisingly good performance of pre-trained LLMs on regression tasks (for example, time-series prediction), with the a
arXiv:2605.13188v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in settings where the available context is incomplete or degraded. We argue that an LLM generat
arXiv:2603.02245v3 Announce Type: replace-cross Abstract: Decoding infant cry causes remains challenging for healthcare monitoring due to short nonstationary signals, limited annotations, and strong d
arXiv:2605.13068v1 Announce Type: new Abstract: Nonlinear inverse problems often trade inexpensive but fragile first-order updates against curvature-aware methods such as Gauss-Newton and Levenberg-Ma
arXiv:2510.19304v3 Announce Type: replace Abstract: Discrete diffusion models offer a promising alternative to autoregressive generation through parallel decoding, but they suffer from a sampling wall
arXiv:2605.12752v1 Announce Type: new Abstract: LoRA is widely adopted for continual fine-tuning of Large Language Models due to its parameter efficiency, modularity across tasks, and compatibility wi
arXiv:2605.13218v1 Announce Type: new Abstract: Cancer is one of the leading causes of death worldwide, making the development of rapid, minimally invasive, label-free and scalable diagnostic strategi
arXiv:2605.13496v1 Announce Type: cross Abstract: Large Language Models (LLMs) have become increasingly prevalent in cloud-based platforms, propelled by the introduction of AI-based consumer and enter