Expand Neurons, Not Parameters
arXiv:2510.04500v2 Announce Type: replace Abstract: This work demonstrates how increasing the number of neurons in a network without increasing its total number of non-zero parameters improves perform
Knowledge catalogue
arXiv:2510.04500v2 Announce Type: replace Abstract: This work demonstrates how increasing the number of neurons in a network without increasing its total number of non-zero parameters improves perform
arXiv:2510.16138v2 Announce Type: replace Abstract: Existing expert merging strategies for Sparse Mixture of Experts (SMoE) typically rely on input-dependent or input-independent averaging of expert p
arXiv:2602.07285v2 Announce Type: replace Abstract: Binary classification based on predicted probabilities (scores) is a fundamental task in supervised machine learning. While thresholding scores is B
arXiv:2506.01467v3 Announce Type: replace Abstract: Graph generative models perform well on small structured data but struggle to scale to large, complex structures. Hierarchical approaches improve sc
arXiv:2509.25906v2 Announce Type: replace Abstract: Federated Learning (FL) often adopts differential privacy (DP) to protect client data, but the added noise required for privacy guarantees can subst
arXiv:2605.31423v1 Announce Type: new Abstract: We introduce universal transformers: fixed transformers that can simulate any transformer in a given class via a suitable input embedding. Analogous to
arXiv:2605.30749v1 Announce Type: new Abstract: Maximum entropy reinforcement learning (MaxEnt-RL) enables robust exploration, yet practical implementations often restrict policies to simple Gaussians
arXiv:2605.31189v1 Announce Type: new Abstract: Tabular prediction in high-stakes domains requires models that are accurate, transparent, and robust to imperfect inputs. We propose FlagGAM, a rule-def
arXiv:2602.02680v2 Announce Type: replace Abstract: The growing scale of deep neural networks, encompassing large language models (LLMs) and vision transformers (ViTs), has made training from scratch
arXiv:2601.22787v2 Announce Type: replace Abstract: Post-training compression is currently divided into two contrasting regimes. On the one hand, fast, data-free, and model-agnostic methods (e.g., NF4
arXiv:2605.31438v1 Announce Type: new Abstract: Time series forecasting often requires learning nonlinear and time-delayed dependencies. A paradigmatic class of forecasting models are nonlinear vector
arXiv:2504.10564v3 Announce Type: replace-cross Abstract: We introduce FLOWR, a novel structure-based framework for the generation and optimization of three-dimensional ligands. FLOWR integrates conti
arXiv:2605.30858v1 Announce Type: new Abstract: Agentic forecasting is important for decision-making in dynamic environments, but it remains challenging because agents must reason from incomplete, tim
arXiv:2405.07836v5 Announce Type: replace Abstract: We introduce Hyper-Trees as a novel framework for modeling time series data using gradient boosted trees. Unlike conventional tree-based approaches
arXiv:2605.31317v1 Announce Type: new Abstract: Machine unlearning aims to remove the influence of selected training examples without full retraining. Standard evaluations often summarize unlearning q
arXiv:2605.31257v1 Announce Type: new Abstract: Fraud detection in payment networks relies on labels generated through heterogeneous and imperfect observation processes, yet existing approaches treat
arXiv:2605.31063v1 Announce Type: cross Abstract: Free energy estimation is a fundamental yet challenging problem, from physics to statistics. Classical approaches rely on thermodynamic transformation
arXiv:2601.19448v2 Announce Type: replace Abstract: Deep Neural Networks remain inherently vulnerable to backdoor attacks. Traditional test-time defenses largely operate under the paradigm of internal
arXiv:2605.30371v1 Announce Type: cross Abstract: We address the issue of global convergence in stochastic continuous optimization. For that purpose, we formulate the Canonical Evolutionary Strategy (
arXiv:2605.31559v1 Announce Type: new Abstract: Learning mappings between infinite-dimensional function spaces, or operator learning, is essential for many machine learning applications. Although tran
arXiv:2605.30374v1 Announce Type: new Abstract: Estimating hip muscle forces and joint moments during gait typically relies on musculoskeletal simulation, which is informative but time-consuming and d
arXiv:2605.31318v1 Announce Type: new Abstract: Modeling an opponent's intent is critical for effective decision-making in non-cooperative, competitive, and general-sum multi-agent reinforcement learn
arXiv:2605.31129v1 Announce Type: new Abstract: Multi-scale modeling has emerged as an effective design principle for time-series forecasting by capturing temporal dynamics at multiple resolutions. As
arXiv:2603.09936v2 Announce Type: replace Abstract: Generative Modeling via Drifting~itep{deng2026drifting} has recently achieved state-of-the-art one-step image generation through a kernel-based drif
arXiv:2605.30453v1 Announce Type: cross Abstract: Generative machine learning has become an essential tool in theoretical and experimental physics, especially in the context of fast surrogates and den
arXiv:2605.30866v1 Announce Type: cross Abstract: Many practically relevant applications of quantum machine learning involve classical data, for which performance depends critically on how inputs are
arXiv:2605.31193v1 Announce Type: new Abstract: Real-world multimodal systems must be robust against low-quality data, such as sensor noise, incomplete multimodal data and conflicting inputs. However,
arXiv:2605.31277v1 Announce Type: cross Abstract: Traditional traffic analysis is being fundamentally challenged by the rapid adoption of encryption, tunnelling, and privacy-preserving protocols, whic
arXiv:2605.31580v1 Announce Type: new Abstract: Transformer-based architectures have advanced sequence modeling in language and vision, yet general-purpose representation learning for heterogeneous mu
arXiv:2601.19966v2 Announce Type: replace-cross Abstract: We introduce ELECTRAFI, a fast, end-to-end differentiable model for predicting periodic charge densities in crystalline materials. ELECTRAFI c
arXiv:2605.30865v1 Announce Type: new Abstract: Continuous glucose monitoring (CGM) provides a dense view of daily metabolic physiology, yet existing generic time-series and CGM-specific foundation mo
arXiv:2511.10868v2 Announce Type: replace Abstract: Training data imbalance poses a major challenge for code LLMs. Most available data heavily over represents raw opensource code while underrepresenti
arXiv:2605.31315v1 Announce Type: new Abstract: We show that contrary to conventional wisdom in the community, graph neural networks (GNNs) are not continuous with respect to all natural modes of grap
arXiv:2605.31485v1 Announce Type: new Abstract: Architecture diagrams are ubiquitous in deep learning, but they are usually only representational: the tensor-program identities they suggest are still
arXiv:2603.08651v2 Announce Type: replace Abstract: We introduce a comprehensive theoretical and algorithmic framework that bridges formal group theory and group entropies with modern machine learning
arXiv:2605.30997v1 Announce Type: cross Abstract: When a learner faces a new task with few samples, it must leverage any available side information. In practice, this often comes in the form of model
arXiv:2605.31000v1 Announce Type: cross Abstract: Training Large Language Models (LLMs) on heterogeneous clusters presents significant challenges for collective communication, as hardware from multipl
arXiv:2602.21340v2 Announce Type: replace Abstract: Representing the past in a compressed, efficient, and informative manner is a central problem for systems trained on sequential data. The HiPPO fram
arXiv:2602.09309v2 Announce Type: replace-cross Abstract: Every generative model for crystalline materials harbors a critical structure size beyond which its outputs become unreliable; we call this th
arXiv:2605.31186v1 Announce Type: new Abstract: Data streams are nowadays among the most frequently analyzed data structures, with the concept drift posing a major challenge encountered by processing
arXiv:2603.11600v2 Announce Type: replace Abstract: Deep reinforcement learning for continuous control often suffers from high variance, low energy efficiency, and poor generalization under distributi
arXiv:2408.16457v5 Announce Type: replace Abstract: Hypergraphs are powerful mathematical structures that can model complex, high-order relationships in various domains, including social networks, bio
arXiv:2601.21645v2 Announce Type: replace Abstract: We investigate the relation between end-to-end equivariance and layerwise equivariance in deep neural networks. We prove the following: For a networ
arXiv:2603.26506v2 Announce Type: replace-cross Abstract: Connectivity structure shapes neural computation, but inferring this structure from population recordings is degenerate: multiple connectivity
arXiv:2605.31413v1 Announce Type: cross Abstract: We establish improved nonasymptotic bounds for Langevin Monte Carlo in the strongly log-concave setting, when the error is measured by the Wasserstein
arXiv:2605.30596v1 Announce Type: new Abstract: Independently trained neural models typically converge to incompatible latent representations, creating a fundamental barrier to highly modular AI syste
arXiv:2605.30615v1 Announce Type: new Abstract: In selective classification, a model predicts the labels of data samples where it is confident, and abstains from predicting labels for samples on which
arXiv:2511.21513v2 Announce Type: replace Abstract: Deploying Transformer models on edge devices is limited by latency and energy budgets. While INT8 quantization effectively accelerates the primary m
arXiv:2604.02969v2 Announce Type: replace-cross Abstract: The natural gradient method is a central tool for statistical optimisation, but its broader application is hindered by the assumption of a Euc
arXiv:2605.30810v1 Announce Type: new Abstract: High-dimensional biomedical data, such as cell-by-gene matrices, are increasingly generated temporally. However, Manifold Learning algorithms, like t-SN
arXiv:2602.09405v2 Announce Type: replace-cross Abstract: We examine the connection between training error and generalization error for arbitrary estimating procedures, working in an overparameterized
arXiv:2605.30741v1 Announce Type: cross Abstract: Epistemic uncertainty quantification (UQ) for deep neural networks (DNNs) is a requirement for safe adoption of AI in mission-critical settings. Sever
arXiv:2605.30622v1 Announce Type: cross Abstract: Open radio access network (O-RAN) architectures enable near real-time, software-driven control of network slicing through programmable xApps deployed
arXiv:2605.30359v1 Announce Type: cross Abstract: Generating high-performance GPU kernels remains challenging due to the need for both correctness and hardware-aware optimization. While large language
arXiv:2603.08721v2 Announce Type: replace-cross Abstract: New AI accelerators with novel instruction set architectures (ISAs) often require developers to manually craft low-level kernels, a time-consu
arXiv:2602.18837v2 Announce Type: replace Abstract: Despite their theoretical advantages, spectral methods based on the graph Fourier transform (GFT) are seldom used in graph neural networks (GNNs) du
arXiv:2510.00419v2 Announce Type: replace Abstract: Zeroth-order optimizers have recently emerged as an attractive approach for fine-tuning large language models (LLMs), as they avoid backpropagation
arXiv:2410.19153v2 Announce Type: replace Abstract: In neuroscience, numerous studies conduct sensory or behavioral experiments under multiple conditions to acquire neural responses in the form of hig
arXiv:2605.30432v1 Announce Type: cross Abstract: Social systems consist of networks of individuals who influence one another through social interactions. Studying how processes evolve on these networ
arXiv:2605.30603v1 Announce Type: cross Abstract: Floating-material transport is influenced by unresolved processes that are often absent from available circulation products. We develop a data-driven