Instance-Adaptive Online Multicalibration
arXiv:2605.09273v1 Announce Type: new Abstract: We study online multicalibration beyond the worst-case. We give a single, efficient algorithm which dynamically interpolates between benign and worst-ca
Knowledge catalogue
arXiv:2605.09273v1 Announce Type: new Abstract: We study online multicalibration beyond the worst-case. We give a single, efficient algorithm which dynamically interpolates between benign and worst-ca
arXiv:2412.01324v4 Announce Type: replace Abstract: This work presents a novel and efficient nonlinear programming framework that tightly integrates hierarchical decision-making with whole-body invers
arXiv:2605.10627v1 Announce Type: cross Abstract: Coreference resolution is typically evaluated using aggregate statistical metrics such as CoNLL-F1, which measure structural overlap between predicted
arXiv:2605.10796v1 Announce Type: new Abstract: Machine learning has become increasingly prevalent in football performance analysis, yet most studies prioritize predictive accuracy while implicitly as
arXiv:2605.10633v1 Announce Type: cross Abstract: Fine-tuning Large Language Models (LLMs) on benign narrow data can sometimes induce broad harmful behaviors, a vulnerability termed emergent misalignm
arXiv:2605.09238v1 Announce Type: cross Abstract: Muon and related norm-constrained matrix optimizers have become central to large-scale learning problems. They are formulated as a linear maximization
arXiv:2605.09439v1 Announce Type: new Abstract: Generative models are powerful tools for sampling from a learned distribution P(Y mid X), and inverse-design methods invert this map to find an input x
arXiv:2605.10684v1 Announce Type: cross Abstract: Data selection studies the problem of identifying high-quality subsets of training data. While some existing works have considered selecting the subse
arXiv:2605.10551v1 Announce Type: new Abstract: Graph Neural Networks (GNNs) have achieved strong results in molecular property prediction, but polymers present distinct challenges: labeled datasets a
arXiv:2605.10178v1 Announce Type: cross Abstract: Adaptive behavior requires the brain to transition between distinct contexts while maintaining representations of prior experience. The ability to rec
arXiv:2605.08587v1 Announce Type: cross Abstract: Long-context language modeling remains central to modern sequence modeling, but the quadratic cost of Transformer attention makes scaling computationa
arXiv:2605.10727v1 Announce Type: new Abstract: We propose kernel-gradient drifting, a one-step generative modeling framework that replaces the fixed Euclidean displacement direction in drifting model
arXiv:2605.09877v1 Announce Type: cross Abstract: We present Key-Value Means ('KVM'), a novel block-recurrence for attention that can accommodate either fixed-size or growing state. Equipping a strong
arXiv:2605.09386v1 Announce Type: cross Abstract: Metric-induced discrete flow matching (MI-DFM) exploits token-latent geometry for discrete generation, but its practical use is limited by two issues:
arXiv:2605.09213v1 Announce Type: cross Abstract: We study causal self-attention dynamics -- a toy model for decoder Transformers -- which we interpret as a non-exchangeable interacting particle syste
arXiv:2605.10253v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) is a widely adopted paradigm for enhancing LLMs in medical applications by incorporating expert multimodal knowle
arXiv:2605.08806v1 Announce Type: new Abstract: Existing 2D-3D lifting human pose estimation methods have achieved strong performance. But the utilization of historical pose representations across net
arXiv:2605.09994v1 Announce Type: cross Abstract: Modern Large Foundation Model (LFM) training has transformed the data pipeline from a static ingestion layer into a dynamic component that must co-evo
arXiv:2605.09060v1 Announce Type: new Abstract: Multilingual vision-language models exhibit systematic performance gaps across languages, but the mechanism remains ambiguous: cross-language divergence
arXiv:2602.09297v2 Announce Type: replace Abstract: Transformers update token representations through multi-head attention and residual connections as X leftarrow X + sum_{i} P^{(i)}XW_{V_i}W_{o_i}, w
arXiv:2605.09204v1 Announce Type: new Abstract: Backpropagation is inherently sequential across depth, creating an O(K)-deep dependency chain that bottlenecks parallel training. While parallel-scan fo
arXiv:2605.08552v1 Announce Type: cross Abstract: Independent Component Analysis (ICA) is a foundational tool for unsupervised representation learning, yet its high-dimensional theory remains largely
arXiv:2605.08209v1 Announce Type: new Abstract: Deep learning methods are widely used under diverse resource constraints, resulting in models of varying sizes, such as the Vision Transformer (ViT) ser
arXiv:2605.09993v1 Announce Type: new Abstract: Graph foundation models (GFMs), pretrained on massive graph data, have transformed graph machine learning by supporting general-purpose reasoning across
arXiv:2605.10855v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have demonstrated remarkable progress in chart understanding, largely driven by supervised fine-tuning (SFT) on increasing
arXiv:2605.08506v1 Announce Type: new Abstract: Robust optimization (RO) provides a principled framework for decision-making under uncertainty, but its performance critically depends on the choice of
arXiv:2511.17994v2 Announce Type: replace Abstract: We study differentially private model training with stochastic gradient descent under learning rate scheduling and correlated noise. Although correl
arXiv:2605.09718v1 Announce Type: cross Abstract: Many systems in physics, engineering, and biology exhibit multiscale stochastic dynamics, where low-dimensional slow variables evolve under the influe
arXiv:2303.05307v2 Announce Type: replace-cross Abstract: We study general-sum, multi-player stochastic games with transferable utility, motivated by settings where agents can use side payments to mak
arXiv:2605.09964v1 Announce Type: new Abstract: Protein-protein interactions (PPIs) are fundamental to cellular function and disease mechanisms. Current learning-based PPI predictors focus on learning
arXiv:2605.08811v1 Announce Type: cross Abstract: This paper investigates the learning theory of Transformer networks for regression tasks on the compact Euclidean domain [0,1]^d and d-dimensional com
arXiv:2601.21410v3 Announce Type: replace-cross Abstract: Large language models (LLMs) encode rich semantic knowledge that can be useful for supervised learning, but their outputs are unreliable as st
arXiv:2605.06394v1 Announce Type: cross Abstract: These lecture notes introduce some topics of classical statistical physics, particularly those that are relevant for neural networks and deep learning
arXiv:2512.23025v2 Announce Type: replace-cross Abstract: Multimodal health sensing offers rich behavioral signals for assessing mental health, yet translating these numerical time-series measurements
arXiv:2509.25742v4 Announce Type: replace Abstract: Graph Contrastive Learning (GCL) has shown strong promise for unsupervised graph representation learning, yet its effectiveness on heterophilic grap
arXiv:2605.09764v1 Announce Type: cross Abstract: LLM-guided evolutionary methods such as AlphaEvolve have proven effective in domains like math, systems research, and algorithmic discovery, but their
arXiv:2503.14434v3 Announce Type: replace-cross Abstract: Automated feature engineering plays a critical role in improving predictive model performance for tabular learning tasks. Traditional automate
arXiv:2605.08850v1 Announce Type: cross Abstract: We design Local LMO - a new projection-free gradient-type method for constrained optimization. The key algorithmic idea is to replace the global linea
arXiv:2605.09830v1 Announce Type: cross Abstract: We present Loom, an outfit recommendation system that combines neural embedding retrieval with structured domain scoring to generate complete, coheren
arXiv:2605.09156v1 Announce Type: cross Abstract: The diachronic evolution from Latin to the Romance languages involved a restructuring of the grammatical gender system from a tripartite configuration
arXiv:2605.09312v1 Announce Type: new Abstract: Neural Radiance Fields (NeRF) achieve high-quality novel-view synthesis, but their long training times and reliance on dense input views limit accessibi
arXiv:2410.10247v3 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have demonstrated exceptional generalization capabilities for downstream tasks. Due to its efficiency, prompt le
arXiv:2605.08996v1 Announce Type: new Abstract: Graph-based accelerators have been widely adopted in symbolic data processing applications such as genomics, cybersecurity, and artificial intelligence.
arXiv:2508.03829v2 Announce Type: replace Abstract: The growing deployment of Large Language Models (LLMs) has raised concerns about their misuse in generating harmful or deceptive content. To address
arXiv:2605.10196v1 Announce Type: new Abstract: High-throughput gene perturbation experiments can test several genetic interventions in parallel, yet experimental budgets remain limited. A central goa
arXiv:2602.07052v2 Announce Type: replace Abstract: Neuronavigation is widely used in biomedical research and interventions to guide the precise placement of instruments around the head to support pro
arXiv:2605.10859v1 Announce Type: new Abstract: Diffusion models dominate image editing, yet their global denoising mechanism entangles edited regions with surrounding context, causing modifications t
arXiv:2605.09236v1 Announce Type: cross Abstract: While digitized corpora have transformed the study of intellectual transmission, current methods rely heavily on lexical text reuse detection, capturi
arXiv:2601.22320v2 Announce Type: replace Abstract: We study continual mean estimation, where data vectors arrive sequentially and the goal is to maintain accurate estimates of the running mean. We ad
arXiv:2605.08863v1 Announce Type: new Abstract: Hallucination detection has become increasingly important for improving the reliability of large language models (LLMs). Recently, hybrid approaches suc
arXiv:2605.10008v1 Announce Type: cross Abstract: Optical readout in low-light imaging is fundamentally limited by measurement noise, including photon shot noise, detector noise, and quantization erro
arXiv:2605.08777v1 Announce Type: cross Abstract: Mode separation, namely how sharply a distribution fragments into barrier-separated clusters, is a fundamental geometric property of densities, diffic
arXiv:2605.10606v1 Announce Type: cross Abstract: Large language models (LLMs) can convincingly imitate human writing styles, yet it remains unclear how much stylistic information is encoded in embedd
arXiv:2509.22196v2 Announce Type: replace Abstract: Disentangled representations seek to recover latent factors of variation underlying observed data, yet their identifiability is still not fully unde
arXiv:2605.08094v1 Announce Type: cross Abstract: Accurate clinical diagnosis requires extensive domain knowledge and complex clinical reasoning capabilities. Although large language models (LLMs) hol
arXiv:2605.10537v1 Announce Type: new Abstract: Memory consolidation, the process by which transient experiences are transformed into stable, structured representations, is a foundational organizing p
arXiv:2605.09270v1 Announce Type: cross Abstract: Supervised Fine-Tuning (SFT) is widely used for task-specific adaptation, yet recent work shows it systematically undermines reasoning generalization.
arXiv:2605.08744v1 Announce Type: cross Abstract: Autoregressive (AR) models can generate high-quality low-poly meshes from point clouds, but they still operate in an all-or-nothing manner: when a loc
arXiv:2602.22508v2 Announce Type: replace Abstract: Large Language Models (LLMs) often produce incorrect answers on multi-hop question answering even when the reasoning trace already contains a correc
arXiv:2605.08701v1 Announce Type: new Abstract: This data paper describes METBRA25Y, a harmonized archive of hourly surface meteorological observations from Brazil derived from public historical recor