Learning Unbiased Permutations via Flow Matching
arXiv:2605.16755v1 Announce Type: cross Abstract: Learning permutations is fundamental to sorting, ranking, and matching, but existing differentiable methods based on entropy-regularized Sinkhorn prod
Knowledge catalogue
arXiv:2605.16755v1 Announce Type: cross Abstract: Learning permutations is fundamental to sorting, ranking, and matching, but existing differentiable methods based on entropy-regularized Sinkhorn prod
arXiv:2605.17779v1 Announce Type: new Abstract: Generative recommendation reformulates recommendation as next-token prediction over discrete semantic identifiers (IDs). A fundamental yet unexplored de
arXiv:2605.16786v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly needed for interactive mobile applications, but high-quality models exceed the limited DRAM available on s
arXiv:2605.17333v1 Announce Type: new Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) typically samples multiple responses per prompt and assigns binary rewards based on individual cor
arXiv:2605.18211v1 Announce Type: cross Abstract: We introduce Graph-Augmented Sequence-to-Sequence (GA-S2S), a novel framework that integrates a T5-small encoder-decoder with a Relational Graph Atten
arXiv:2605.12987v2 Announce Type: replace Abstract: BACKGROUND: Coding Motivational Interviewing (MI) sessions is essential for understanding client behaviors and predicting outcomes, but it requires
arXiv:2506.21499v2 Announce Type: replace-cross Abstract: Ultrasound Coherent Plane-Wave Compounding (CPWC) enhances image contrast by combining echoes from multiple steered transmissions. While incre
arXiv:2603.16947v2 Announce Type: replace-cross Abstract: Although vision-language navigation (VLN) has progressed rapidly, zero-shot VLN in continuous environments (VLN-CE) remains highly challenging
arXiv:2509.01629v3 Announce Type: replace-cross Abstract: We study the design of interpolation schedules in flow and diffusion-based generative models from both statistical and numerical perspectives.
arXiv:2605.17287v1 Announce Type: new Abstract: Driver gaze estimation serves as a fundamental metric for evaluating driver attentiveness in modern monitoring systems. Beyond being vulnerable to sudde
arXiv:2605.17260v1 Announce Type: new Abstract: The fundamental challenge in scaling Video Large Language Models (Video LLMs) to long-form video lies in managing the explosion of visual-token context
arXiv:2503.02574v2 Announce Type: replace-cross Abstract: In this paper, we argue that current safety alignment research efforts for large language models are hindered by many intertwined sources of n
arXiv:2503.14800v3 Announce Type: replace-cross Abstract: Effective long-term memory management is crucial for language models handling extended contexts. We introduce the Enhanced Ranked Memory Augme
arXiv:2605.17888v1 Announce Type: cross Abstract: Long-horizon prediction of three-dimensional (3D) wall-bounded turbulence with machine-learning methods remains a challenging task, due to the rapid a
arXiv:2605.17603v1 Announce Type: cross Abstract: High-resolution precipitation information is essential for climate impact assessment, yet global climate models remain too coarse to resolve key small
arXiv:2605.16374v1 Announce Type: cross Abstract: Continual learning studies how models can adapt to new tasks while retaining previously acquired knowledge. Although a broad spectrum of methods has b
arXiv:2512.01030v3 Announce Type: replace Abstract: Recovering pixel-wise geometric properties from a single image is fundamentally ill-posed due to appearance ambiguity and non-injective mappings bet
arXiv:2601.14330v2 Announce Type: replace Abstract: Concept erasure aims to suppress sensitive content in diffusion models, but recent studies show that erased concepts can still be reawakened, reveal
arXiv:2605.16365v1 Announce Type: new Abstract: Early identification of individuals at elevated risk of Chlamydia trachomatis infection may enable optimal use of molecular testing in resource-aware sc
arXiv:2605.17120v1 Announce Type: new Abstract: arly identification of motor impairment in infancy relies on expert visual assessment of spontaneous movement, motivating the development of automated,
arXiv:2605.17640v1 Announce Type: cross Abstract: Retrieval-augmented generation from videos requires systems to retrieve relevant audiovisual evidence from large corpora and synthesize it into cohere
arXiv:2605.06017v2 Announce Type: replace Abstract: Sequence-level evaluations in autoregressive Large Language Models (LLMs) rely on highly dependent token generation. Establishing tight concentratio
arXiv:2605.17091v1 Announce Type: new Abstract: Scientific forecasting typically relies on direct state prediction, an approach that grows brittle under data scarcity, extended horizons, non-stationar
arXiv:2605.16468v1 Announce Type: cross Abstract: A central goal in understanding human vision is to uncover the visual features that drive neuronal activity. A growing body of work has used artificia
arXiv:2605.16639v1 Announce Type: new Abstract: Multimodal clinical prediction faces three challenges: multiple foundation models (FMs) with complementary strengths per modality, pervasive missing mod
arXiv:2605.17365v1 Announce Type: new Abstract: Different from traditional text-to-image retrieval tasks, chat-based image retrieval allows the human-interactive system to iteratively clarify and refi
arXiv:2506.15588v2 Announce Type: replace Abstract: Differential privacy (DP) protects sensitive data during neural network training, but standard methods like DP-Adam suffer from high memory overhead
arXiv:2605.17539v1 Announce Type: new Abstract: Combinatorial optimization (CO) underlies decision-making from logistics to chip design, where infeasible solutions are operationally unusable and small
arXiv:2507.22057v2 Announce Type: replace Abstract: Difficult few-shot image recognition has significant application prospects, yet remaining the substantial technical gaps with the conventional large
arXiv:2605.16864v1 Announce Type: cross Abstract: Although large-scale visual foundation models (VFMs) achieve remarkable performance in semantic understanding, they still underperform in instance-awa
arXiv:2605.16464v1 Announce Type: cross Abstract: Brain tumors exhibit high heterogeneity in morphology and multimodal contrast, making manual slice-by-slice de lineation time-consuming and experience
arXiv:2505.02621v2 Announce Type: replace Abstract: The mean-field Langevin dynamics (MFLD) minimizes an entropy-regularized nonlinear convex functional on the Wasserstein space over R^d, and has gain
arXiv:2507.06384v2 Announce Type: replace-cross Abstract: Objective: Latent diffusion models (LDMs) could mitigate data scarcity challenges affecting machine learning development for medical image int
arXiv:2402.15058v3 Announce Type: replace-cross Abstract: We combine standard persistent homology with image persistent homology to define a novel way of characterizing shapes and interactions between
arXiv:2605.17635v1 Announce Type: cross Abstract: A fast simulation of the detector response is a vital task in high-energy physics (HEP). Traditional Monte-Carlo methods form the backbone of modern p
arXiv:2603.19538v2 Announce Type: replace Abstract: Monocular 3D object understanding has largely been cast as a 2D RoI-to-3D box lifting problem. However, emerging downstream applications require ima
arXiv:2605.18483v1 Announce Type: cross Abstract: Time series classification (TSC) of biological signals has progressed from handcrafted, modality-specific approaches to deep architectures capable of
arXiv:2605.17252v1 Announce Type: new Abstract: Stereoscopic 3D displays adopt a binocular depth cue to provide depth perception. However, users should be equipped with expensive special devices to ap
arXiv:2605.16932v1 Announce Type: new Abstract: Robots deployed in unstructured human environments must frequently execute long-horizon missions, such as find the mug, then the chair, then the printer
arXiv:2502.05462v2 Announce Type: replace Abstract: We propose a real-time implementable motion planning framework for cooperative object transportation by nonholonomic mobile manipulator robots (MMRs
arXiv:2605.16456v1 Announce Type: new Abstract: Understanding how objects relate to each other in space is fundamental to scene understanding, yet most contrastive pre-training approaches only model p
arXiv:2604.00919v2 Announce Type: replace-cross Abstract: Energy-based models provide a natural bridge between statistical physics and machine learning by representing data through structured energy l
arXiv:2605.17624v1 Announce Type: cross Abstract: We investigate the potential of invariant and equivariant semi-supervised learning for addressing the challenges of training multi-task models on part
arXiv:2605.17761v1 Announce Type: cross Abstract: Insider threats often reveal early anomalies through disruptions in behavioral statistics-such as altered recurrence patterns or short-versus long-ter
arXiv:2505.11143v2 Announce Type: replace-cross Abstract: Sparse linear regression is a fundamental tool in data analysis. However, traditional approaches often fall short when covariates exhibit stru
arXiv:2510.16814v2 Announce Type: replace-cross Abstract: Archaeological predictive modelling estimates where undiscovered sites are likely to occur by combining known locations with environmental, cu
arXiv:2605.16805v1 Announce Type: new Abstract: LiDARs are widely used for 3D depth reconstruction, but their performance is often limited by inherent hardware constraints that impose trade-offs betwe
arXiv:2510.13068v4 Announce Type: replace-cross Abstract: Biosignals such as electroencephalography (EEG), electrocardiography (ECG), and electromyography (EMG) encode physiological activity across mu
arXiv:2605.18035v1 Announce Type: new Abstract: Hard-thresholding is an important type of algorithm in machine learning that is used to solve ell_0 constrained optimization problems. However, the true
arXiv:2605.16893v1 Announce Type: new Abstract: Recent studies introduce conditional memory modules that decouple knowledge storage from neural computation, enabling more direct knowledge access. Comp
arXiv:2605.16514v1 Announce Type: cross Abstract: Understanding why some sequential planning problems are harder than others requires models that go beyond average performance. They should capture the
arXiv:2605.17390v1 Announce Type: cross Abstract: Context. Metamorphic Testing is recognised in IEEE/ISO software-testing standards and increasingly recommended for AI systems, but its progress is bot
arXiv:2605.18238v1 Announce Type: new Abstract: Digital entities such as AI agents and humanoid robots increasingly operate alongside real humans, yet their identity infrastructure is based on credent
arXiv:2605.16423v1 Announce Type: new Abstract: Network quantization has emerged as one of the most practical model compression techniques, which significantly reduces a model's memory and compute con
arXiv:2605.17340v1 Announce Type: new Abstract: Time series foundation models rely on large-scale pretraining over diverse datasets across domains, yet their heterogeneity in temporal patterns could h
arXiv:2605.17488v1 Announce Type: new Abstract: The landscape of joint audio and video generation has been fundamentally transformed by the advent of powerful foundation models. Despite these strides,
arXiv:2605.18041v1 Announce Type: new Abstract: Omnimodal large language models (OmniLLMs) have recently gained increasing attention for unified audio-video understanding. However, processing long mul
arXiv:2605.16998v1 Announce Type: cross Abstract: The Quantum Fourier Transform (QFT) is required by hidden subgroup problem (HSP) algorithms, including Shor's algorithm for factoring. The circuit dep
arXiv:2605.17483v1 Announce Type: new Abstract: Facial Expression Recognition faces two core challenges. The first is class imbalance in public datasets, which skews the learning process and weakens g
arXiv:2605.17678v1 Announce Type: cross Abstract: In this paper, we derive rates of convergence in the high-dimensional central limit theorem for Polyak--Ruppert averaged iterates generated by entropy