(Mis)generalization of Helpful-only Fine-tuning
arXiv:2606.04413v1 Announce Type: new Abstract: Helpful-only models, that is, models that are trained to always follow user intent, are valuable for dangerous capability evaluations and other areas of
Knowledge catalogue
arXiv:2606.04413v1 Announce Type: new Abstract: Helpful-only models, that is, models that are trained to always follow user intent, are valuable for dangerous capability evaluations and other areas of
arXiv:2606.04499v1 Announce Type: cross Abstract: Cancer care requires a longitudinal approach in which treatments are planned and delivered over time according to the needs of each individual patient
arXiv:2606.04757v1 Announce Type: cross Abstract: We study decentralized stochastic smooth convex optimization, where M workers minimize an average objective using local stochastic gradients and neigh
arXiv:2606.04324v1 Announce Type: new Abstract: One of the primary challenges in Bayesian inference on the parameters of a diffusion model from discrete observations is the unavailability of an analyt
arXiv:2506.23546v2 Announce Type: replace-cross Abstract: Fixed points of recurrent neural networks can be leveraged to store and generate information. These fixed points can be captured by the Boltzm
arXiv:2606.04994v1 Announce Type: new Abstract: Accurate computational prediction of T cell receptor (TCR) antigen specificity would transform the study of T cell biology and enable scalable immune en
arXiv:2606.04957v1 Announce Type: cross Abstract: System-generated logs underpin security monitoring, yet their rigid template-based format hinders both automated analysis and human comprehension. We
arXiv:2606.04265v1 Announce Type: cross Abstract: The Schrodinger Bridge Problem constructs a stochastic process that connects an initial distribution to a terminal distribution with minimum energy. T
arXiv:2606.04028v1 Announce Type: new Abstract: The IEEE P3109 draft standard defines a parameterized family of binary floating-point formats and associated operations, with a focus on facilitating ma
arXiv:2606.04305v1 Announce Type: new Abstract: We study online learning with an additional offline dataset in the stochastic linear bandit setting. Although this problem arises frequently in practice
arXiv:2601.21868v2 Announce Type: replace-cross Abstract: Understanding the stability and long-time behavior of generative models is a fundamental problem in modern machine learning. This paper provid
arXiv:2606.04451v1 Announce Type: new Abstract: Neighbor embedding algorithms reveal correlations in high-dimensional data by constructing an equivalent graph representation in a lower-dimensional spa
arXiv:2602.01083v2 Announce Type: replace Abstract: Weight-space learning studies neural architectures that operate directly on the parameters of other neural networks. Motivated by the growing availa
arXiv:2502.00470v3 Announce Type: replace-cross Abstract: Distributed empirical risk minimization (ERM) is often studied through two influential yet seemingly separate families of methods: CoCoA-type
arXiv:2601.07144v3 Announce Type: replace-cross Abstract: Ensuring fairness in matching algorithms is a key challenge in allocating scarce resources and positions. Focusing on Optimal Transport (OT),
arXiv:2604.00915v2 Announce Type: replace Abstract: Estimation of heterogeneous long-term treatment effects (HLTEs) is relevant for personalized decision-making in marketing, economics, and medicine,
arXiv:2509.22454v2 Announce Type: replace Abstract: Electrostatic generative models such as PFGM++ have recently emerged as a powerful framework, achieving competitive performance in image synthesis.
arXiv:2602.19799v2 Announce Type: replace-cross Abstract: Despite recent algorithmic advances, we still lack principled ways to leverage the well-documented rescaling symmetries in ReLU neural network
arXiv:2606.04290v1 Announce Type: new Abstract: Hybrid models that combine physics-based and data-driven components have shown strong potential for achieving accuracy and interpretability in control a
arXiv:2606.04335v1 Announce Type: new Abstract: The framework of robust Markov decision processes (RMDPs) allows the design of reinforcement learning agents that satisfy performance guarantees under w
arXiv:2606.04834v1 Announce Type: new Abstract: Minimum Description Length (MDL) formalizes the principle of Occam's razor by optimizing the total description length: L(model)+L(data | model). For seq
arXiv:2606.05129v1 Announce Type: cross Abstract: Preserving data privacy is an important topic in structural data management and data mining. However, the issue of privacy leakage in distributed caus
arXiv:2606.04866v1 Announce Type: new Abstract: Large-scale hyperparameter optimization (HPO) in automated machine learning (AutoML) consumes substantial computational resources, raising growing conce
arXiv:2606.04031v1 Announce Type: new Abstract: Coupled gradient descent--where the update of one parameter block depends on another--underlies bilevel optimization, two-time-scale stochastic approxim
arXiv:2606.04689v1 Announce Type: cross Abstract: Scene Graph Generation (SGG) requires relational reasoning over objects and their interactions, but performance is often limited by severe long-tail p
arXiv:2511.21035v2 Announce Type: replace Abstract: Holography offers significant potential for AR/VR applications. However, its adoption is limited by the high demand for data compression. Existing d
arXiv:2604.01161v2 Announce Type: replace Abstract: Large language models (LLMs) exhibiting test-time scaling behavior, such as extended reasoning traces and self-verification, have demonstrated remar
arXiv:2606.04822v1 Announce Type: new Abstract: Causal modeling of physical temporal phenomena must handle interventions that act along trajectories, nonstationary induced laws, path-dependent effects
arXiv:2606.04582v1 Announce Type: cross Abstract: Real-time monitoring of the temperature distribution within components and sub-structures is a challenging topic in many systems due to restrictions o
arXiv:2004.10846v5 Announce Type: replace-cross Abstract: Problem definition: Traditionally, New York City's top 8 public schools have selected candidates solely based on their scores in the Specializ
arXiv:2606.04380v1 Announce Type: cross Abstract: Forecast reconciliation usually starts from a fixed measurement system and asks how forecasts should be projected onto a coherent space. We ask a diff
arXiv:2606.05109v1 Announce Type: new Abstract: To leverage the full potential of multimodal data, we need representations that go beyond the state-of-the-art alignment and fusion approaches and explo
arXiv:2606.04210v1 Announce Type: cross Abstract: Randomized smoothing (RS) certifies robustness in the vector space where Gaussian noise is added. In audio classification, this space is often not uni
arXiv:2606.04576v1 Announce Type: cross Abstract: Learning Value-at-Risk (VaR) and Expected Shortfall (ES) is important for managing financial risks effectively. Existing approaches with limited param
arXiv:2606.04857v1 Announce Type: new Abstract: Standard IMVC evaluation retrains separate models for different missing-data configurations. We show that this paradigm obscures a fundamental vulnerabi
arXiv:2506.06178v3 Announce Type: replace Abstract: Policy gradient (PG) methods are a class of effective reinforcement learning algorithms, particularly when dealing with continuous control problems.
arXiv:2606.04384v1 Announce Type: new Abstract: Machine learning's reliance on sensitive data necessitates privacy-preserving techniques like Differentially Private Stochastic Gradient Descent (DPSGD)
arXiv:2606.05070v1 Announce Type: new Abstract: Train delay prediction is an important problem for both passengers and railway operators, yet progress in the field remains difficult to assess due to t
arXiv:2606.04272v1 Announce Type: new Abstract: The standard LLM training pipeline applies reinforcement learning (RL) only after pre-training and supervised fine-tuning (SFT). We question this status
arXiv:2606.04066v1 Announce Type: cross Abstract: Understanding how structural connections are associated with tau propagation in Alzheimer's disease (AD) remains a central open question, yet existing
arXiv:2606.04444v1 Announce Type: cross Abstract: Existing datasets cannot support large-scale learning in multi-agent, multi-sensor, or multi-domain autonomy, where diversity and coordination are ess
arXiv:2505.21331v2 Announce Type: replace-cross Abstract: In content moderation for social media platforms, the cost of delaying the review of a content is proportional to its view trajectory, which f
arXiv:2606.04036v1 Announce Type: new Abstract: On-policy self-distillation, where a language model conditions on privileged context to supervise its own generations, is a promising source of dense su
arXiv:2606.04929v1 Announce Type: new Abstract: LLM post-training proceeds through multiple stages, e.g., supervised fine-tuning (SFT) followed by reinforcement learning from human feedback (RLHF) or
arXiv:2602.01027v2 Announce Type: replace Abstract: Mixed-precision quantization is a promising approach for compressing large language models under tight memory budgets. However, existing mixed-preci
arXiv:2606.04390v1 Announce Type: new Abstract: We find the asymptotic ratio between the storage capacities when enforcing real pre-activations in a complex hypothesis class as opposed to complex ones
arXiv:2602.20651v3 Announce Type: replace Abstract: In modern applications such as ECG monitoring, neuroimaging, wearable sensing, and industrial equipment diagnostics, complex and continuously struct
arXiv:2606.04020v1 Announce Type: cross Abstract: Splice-mediated drug resistance occurs in up to 40% of patients on targeted kinase inhibitors, yet state-of-the-art druggability tools operate on sing
arXiv:2606.04000v1 Announce Type: cross Abstract: We present a probabilistic modeling framework for incorporating small-scale spatial heterogeneity into macroscopic descriptions of material behavior f
arXiv:2606.04945v1 Announce Type: new Abstract: Diffusion large language models (DLLMs) have recently emerged as a promising alternative to autoregressive LLMs by generating text through iterative mas
arXiv:2606.04135v1 Announce Type: new Abstract: Time series forecasting relies on historical patterns, but real-world series often exhibit non-stationarity and regime shifts that challenge fully param
arXiv:2606.04100v1 Announce Type: new Abstract: Machine learning interatomic potentials (MLIPs) enable efficient and accurate atomistic simulations but depend critically on the quality and diversity o
arXiv:2606.04021v1 Announce Type: cross Abstract: Proteolysis-targeting chimeras (PROTACs) can selectively degrade disease-causing proteins, yet predicting which targets are amenable to degradation re
arXiv:2606.04564v1 Announce Type: new Abstract: Tabular foundation models (TFMs) have made rapid progress in standard classification and regression, but time-to-event survival prediction tasks have re
arXiv:2601.04051v3 Announce Type: replace Abstract: Symbolic regression aims to find symbolic expressions that describe datasets. Due to its inherent interpretability, symbolic regression (SR) is a po
arXiv:2606.04401v1 Announce Type: new Abstract: The capabilities of large language models (LLMs) significantly depend on training data drawn from various domains. Optimizing domain-specific mixture ra
arXiv:2606.04678v1 Announce Type: new Abstract: End-to-end ASR systems typically use fixed-depth acoustic encoders at inference, making it difficult to trade additional test-time computation for impro
arXiv:2606.04314v1 Announce Type: new Abstract: As neural networks are increasingly deployed in safety-critical domains, testing is essential to evaluate and improve their reliability. Existing testin
arXiv:2602.11406v2 Announce Type: replace-cross Abstract: We consider an online learning problem in environments with multiple change points. In contrast to the single change point problem that is wid
arXiv:2606.04423v1 Announce Type: new Abstract: We show every multi-group learner in the transductive setting may incur a multiplicative penalty in its error rate on some group relative to the error r