Dimension-Free Saddle-Point Escape in Muon
arXiv:2605.09331v1 Announce Type: new Abstract: Modern Large Language Model (LLM) training is fundamentally bottlenecked by pathologically flat saddle points in extreme high-dimensional landscapes. Mo
Knowledge catalogue
arXiv:2605.09331v1 Announce Type: new Abstract: Modern Large Language Model (LLM) training is fundamentally bottlenecked by pathologically flat saddle points in extreme high-dimensional landscapes. Mo
arXiv:2605.08249v1 Announce Type: new Abstract: Frozen vision foundation models do not merely extract features; they organize images through a learned coordinate system. We ask whether that coordinate
arXiv:2605.05284v2 Announce Type: replace-cross Abstract: Evolutionary computation has long promised to deliver both high-performance optimization tools as well as rigorous scientific simulations of D
arXiv:2602.13759v2 Announce Type: replace Abstract: We study eigendecomposition on SO(n) under streaming observations C_k = C_{sig} + sigma_k^2 I + E_k, where the isotropic background sigma_k^2 I may
arXiv:2605.08882v1 Announce Type: new Abstract: Flow Matching has recently emerged as a popular class of generative models for simulating a target distribution mu_1 from samples drawn from a source di
arXiv:2605.09302v1 Announce Type: cross Abstract: We study posterior sampling for inverse problems in discrete state spaces using discrete diffusion models as generative priors. While continuous diffu
arXiv:2605.09881v1 Announce Type: cross Abstract: Mechanistic interpretability seeks to reverse engineer a trained neural network by identifying the minimal subset of internal components. We perform a
arXiv:2605.08104v1 Announce Type: new Abstract: This paper explores the application of the Soft Actor-Critic (SAC) algorithm within a Distributional Reinforcement Learning setting and introduces an im
arXiv:2605.09693v1 Announce Type: cross Abstract: Yes. We find that large multimodal models develop mental imagery when solving spatial puzzles, and they do imagine sheep when solving sheep puzzles. W
arXiv:2605.08199v1 Announce Type: cross Abstract: Cardiovascular disease remains the leading cause of death globally, underscoring the need for effective, accessible monitoring solutions, particularly
arXiv:2605.09514v1 Announce Type: new Abstract: Unobserved confounding prevents standard covariate adjustment from identifying causal response functions in observational studies. Proxy causal learning
arXiv:2605.09347v1 Announce Type: new Abstract: Discrete variables are common in many applications, such as probabilistic reasoning, planning and explainable AI. When symbolic reasoning techniques are
arXiv:2605.09566v1 Announce Type: new Abstract: Recent Deep Unfolding Networks (DUNs) have significantly advanced Compressive Sensing (CS) by integrating iterative optimization with deep networks. How
arXiv:2604.05064v2 Announce Type: replace-cross Abstract: Synthetic data is essential for training foundation models for time series (FMTS), but most generators assume static correlations, and are typ
arXiv:2605.09098v1 Announce Type: new Abstract: We propose Dynamic Meta-Metrics (DMM), a framework for machine translation evaluation that learns source-sentence conditioned combinations of existing m
arXiv:2605.10185v1 Announce Type: cross Abstract: Ghost imaging reconstructs spatial information from a single-pixel bucket detector by correlating structured illumination patterns with scalar intensi
arXiv:2605.10360v1 Announce Type: new Abstract: While novel view synthesis (NVS) for dynamic scenes has seen significant progress, reconstructing temporally consistent geometric surfaces remains a cha
arXiv:2605.10050v1 Announce Type: new Abstract: Long-form video understanding remains challenging for Video Large Language Models (VideoLLMs), as the dense frame sampling introduces massive visual tok
arXiv:2605.09603v1 Announce Type: new Abstract: Masked diffusion language models enable parallel token generation and offer improved decoding efficiency over autoregressive models. However, their perf
arXiv:2410.01656v2 Announce Type: replace-cross Abstract: We study the estimation of distributional parameters when samples are shown only if they fall in some unknown set S subseteq R^d. Kontonis, Tz
arXiv:2605.08606v1 Announce Type: new Abstract: Egocentric human mesh recovery (HMR) from monocular head-mounted cameras is increasingly important for AR/VR applications, but remains challenging due t
arXiv:2605.10938v1 Announce Type: cross Abstract: Diffusion and flow-based models have become the de facto approaches for generating continuous data, e.g., in domains such as images and videos. Their
arXiv:2507.06658v2 Announce Type: replace-cross Abstract: Theories of democratic stability, populism, and party-system crisis often point to a form of polarization that comparative research rarely mea
arXiv:2605.10790v1 Announce Type: new Abstract: Diffusion models have achieved remarkable success, yet their training remains inefficient due to a severe optimization bottleneck, which we term Represe
arXiv:2605.08377v1 Announce Type: new Abstract: In many practical applications it is important to build symmetries into neural network architectures. Consider the important case of permutation symmetr
arXiv:2605.08360v1 Announce Type: new Abstract: Modern AI is opening the door to collective decision-making in which participants express their views as free-form text rather than voting on a fixed se
arXiv:2602.01194v2 Announce Type: replace Abstract: Long-term weather forecasting is critical for socioeconomic planning and disaster preparedness. While recent approaches employ finetuning to extend
arXiv:2605.06663v2 Announce Type: replace Abstract: Large language models are typically deployed as monolithic systems, requiring the full model even when applications need only a narrow subset of cap
arXiv:2605.09509v1 Announce Type: cross Abstract: The problem of predicting unobserved entries in a binary matrix, known as 1-bit matrix completion, has found diverse applications in fields such as re
arXiv:2605.10198v1 Announce Type: cross Abstract: Erasing specific concepts from text-to-image diffusion models is essential for avoiding the generation of copyrighted and explicit content. Closed-for
arXiv:2605.09495v1 Announce Type: cross Abstract: Machine learning-based simulators offer the potential to model the dynamics of complex systems more efficiently than classical approaches, while retai
arXiv:2605.08645v1 Announce Type: cross Abstract: Energy-based models (EBMs) provide a powerful and flexible way of learning a joint probability distribution over data by constructing an energy surfac
arXiv:2605.08910v1 Announce Type: cross Abstract: The new wave of adversarial attacks that utilize gradient-related vulnerabilities in neural network-based classifiers makes Network Intrusion Detectio
arXiv:2601.15065v2 Announce Type: replace Abstract: CLIP-based foreground-background (FG-BG) decomposition methods have demonstrated remarkable effectiveness in improving few-shot out-of-distribution
arXiv:2605.09745v1 Announce Type: cross Abstract: Large language models (LLMs) achieve remarkable generative performance, yet their output quality is dependent on the decoding strategy. While sampling
arXiv:2509.18484v2 Announce Type: replace-cross Abstract: Estimating causal effects on networks is challenging because treatments may affect both treated units and their neighbors, while network homop
arXiv:2605.09429v1 Announce Type: cross Abstract: Are low-attention visual tokens truly redundant in vision-language reasoning? Existing pruning methods often assume so, ranking visual tokens by shall
arXiv:2605.08549v1 Announce Type: new Abstract: Conversational AI is increasingly personalized around users' preferences, histories, goals, and knowledge, but much less around how users interpret and
arXiv:2605.09042v1 Announce Type: new Abstract: Evaluating pragmatic reasoning in large language models (LLMs) remains challenging because model behavior can vary depending on evaluation methods. Prev
arXiv:2605.10402v1 Announce Type: cross Abstract: A finite presentation of a finite group is called `just finite' if removing any relation from R results in a presentation for an infinite group. It ha
arXiv:2605.09935v1 Announce Type: new Abstract: With the rapid development of deep generative models, forged facial images are massively exploited for illegal activities. Although existing synthetic f
arXiv:2605.09924v1 Announce Type: new Abstract: Recent advancements in Neural Machine Translation (NMT) have significantly improved translation quality. However, the increasing size and complexity of
arXiv:2605.10663v1 Announce Type: new Abstract: Experience-driven self-evolving agents aim to overcome the static nature of large language models by distilling reusable experience from past interactio
arXiv:2605.10613v1 Announce Type: cross Abstract: We introduce a technique that enables Neural-ODEs to approximate arbitrary velocity fields with a priori planted fixed-points. Specifically, a recipe
arXiv:2605.10680v1 Announce Type: new Abstract: This paper proposes a paradigm shift linking machine unlearning directly to the structure of the data distributions rather than a mere update of the neu
arXiv:2605.09369v1 Announce Type: new Abstract: Knowledge Tracing (KT) models students' knowledge states based on learning interactions to predict performance. While deep learning-based KT models have
arXiv:2407.07639v2 Announce Type: replace-cross Abstract: Similarity search is a fundamental task for exploiting information in various applications dealing with graph data, such as citation networks
arXiv:2509.20599v2 Announce Type: replace Abstract: Backpropagation through (neural) SDE solvers is traditionally approached in two ways: discretise-then-optimise, which offers accurate gradients but
arXiv:2511.18374v2 Announce Type: replace Abstract: We derive a computable closed-form upper bound on the Hausdorff distance between a truncated minimal robust positively invariant (mRPI) set and its
arXiv:2604.06720v2 Announce Type: replace Abstract: We present DeSOPE, a large-scale dataset for 6DoF deformed objects. Most 6D object pose methods assume rigid or articulated objects, an assumption t
arXiv:2605.08398v1 Announce Type: cross Abstract: In this work, we show that Latent Flow-Matching (LFM) models are robust to different types of perturbations, including data reduction and model capaci
arXiv:2605.10318v1 Announce Type: new Abstract: Large language models (LLMs) allow users to query databases using natural language by translating questions into executable queries. Despite strong prog
arXiv:2605.10045v1 Announce Type: new Abstract: Visual Autoregressive (VAR) models have emerged as a strong alternative to diffusion for image synthesis, yet their fixed training resolution prevents d
arXiv:2605.10795v1 Announce Type: cross Abstract: Large language models demonstrate remarkable ability in factual recall, yet the fundamental limits of storing and retrieving input--output association
arXiv:2605.10330v1 Announce Type: cross Abstract: We propose a novel adaptive Mixture-of-Experts (MoE) framework for time series forecasting that enhances expert specialization by incorporating expert
arXiv:2601.22204v2 Announce Type: replace Abstract: Federated learning (FL) encounters substantial challenges due to heterogeneity, leading to gradient noise, client drift, and partial client particip
arXiv:2605.09428v1 Announce Type: new Abstract: Graph-level anomaly detection (GLAD) is crucial for ensuring the reliability of graph-driven applications by identifying abnormal graphs that deviate fr
arXiv:2602.04093v2 Announce Type: replace Abstract: Concept-based Models (CMs) enhance interpretability in deep learning by grounding predictions in human-understandable concepts. However, concept ann
arXiv:2605.10082v1 Announce Type: new Abstract: Large language models (LLMs) exhibit strong reasoning capabilities when guided by high-quality demonstrations, yet such data is often distributed across
arXiv:2605.09750v1 Announce Type: new Abstract: This article presents a novel approach to keyframe detection in ultrasound videos, with a particular focus on fetal brain imaging. The proposed model is