Muown: Row-Norm Control for Muon Optimization
arXiv:2605.10797v1 Announce Type: new Abstract: Muon has emerged as a strong competitor to AdamW for language model pre-training, yet its behavior at scale is sensitive to weight decay. Recent work ha
Knowledge catalogue
arXiv:2605.10797v1 Announce Type: new Abstract: Muon has emerged as a strong competitor to AdamW for language model pre-training, yet its behavior at scale is sensitive to weight decay. Recent work ha
arXiv:2507.14958v2 Announce Type: replace Abstract: Current models have achieved impressive performance on reasoning-intensive tasks, yet optimizing their reasoning efficiency remains an open challeng
arXiv:2511.07833v3 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become a standard recipe for post-training LLMs on reasoning tasks, with Group Relat
arXiv:2605.10026v1 Announce Type: new Abstract: With the advancement of autonomous driving, numerous annotated multi-modality datasets have become available. This presents an opportunity to develop do
arXiv:2605.09349v1 Announce Type: cross Abstract: We consider a mutual information (MI) regularized version of optimal density control of a discrete-time linear system. MI optimal control has been pro
arXiv:2605.09672v1 Announce Type: new Abstract: State-of-the-art 6-DoF grasp generators excel on tabletop benchmarks with overhead cameras but struggle in frontal grasping scenarios on low-cost manipu
arXiv:2605.09918v1 Announce Type: cross Abstract: Reconciling platform revenue with user experience in LLM advertising motivates a data-centric foundation. We introduce NaiAD, the first comprehensive
arXiv:2605.10210v1 Announce Type: cross Abstract: Terrain segmentation is a fundamental capability for autonomous mobile robots operating in unstructured outdoor environments. However, state-of-the-ar
arXiv:2605.10813v1 Announce Type: new Abstract: LLM-powered multi-agent systems can now automate the full research pipeline from ideation to paper writing, but a fundamental question remains: automati
arXiv:2605.08503v1 Announce Type: new Abstract: Interactive narrative tasks require LLMs to sustain a coherent, evolving story while adapting to a user over multiple turns. However, suitable benchmark
arXiv:2605.08742v1 Announce Type: cross Abstract: This study proposes a quantitative framework for profiling LLM dispositions as stable, model-specific regularities in output under repeated, controlle
arXiv:2605.10671v1 Announce Type: new Abstract: In this work, we show that natural policy gradient, a core algorithm in reinforcement learning, admits an exact formulation as a smoothed and averaged f
arXiv:2605.09863v1 Announce Type: cross Abstract: Production LLM coding agents drift over long sessions: they forget user-specified constraints, slip into mistakes the user already flagged, and confab
arXiv:2605.09176v1 Announce Type: cross Abstract: Training large language models requires optimization algorithms that are not only statistically effective, but also computationally and memory efficie
arXiv:2605.10639v1 Announce Type: new Abstract: The rapid adoption of LLMs in both research and industry highlights the challenges of deploying them safely and reveals a gap in the systematic evaluati
arXiv:2605.10065v1 Announce Type: cross Abstract: Controlling Large Language Models (LLMs) to prevent the generation of undesirable content, such as profanity and personally identifiable information (
arXiv:2605.09363v1 Announce Type: new Abstract: Last-iterate convergence of learning dynamics in games has attracted significant recent attention. In two-player zero-sum games with bandit feedback, wh
arXiv:2605.10299v1 Announce Type: new Abstract: This paper studies kernelized bandits (also known as Gaussian process bandits) in an adversarial environment, where the reward functions in a known repr
arXiv:2605.09778v1 Announce Type: cross Abstract: Evaluating softmax attention over a fixed long context requires reading every cached key-value pair for each new query token. For a given context (a b
arXiv:2510.05635v2 Announce Type: replace-cross Abstract: Test-Time Adaptation (TTA) methods are often computationally expensive, require a large amount of data for effective adaptation, or are brittl
arXiv:2601.23252v2 Announce Type: replace-cross Abstract: Model comparison and calibrated uncertainty quantification often require integrating over parameters, but scalable inference can be challengin
arXiv:2605.09886v1 Announce Type: new Abstract: Generative driving world models rely on compact latent state representations that must be efficiently transmitted and synchronized across distributed co
arXiv:2605.10877v1 Announce Type: new Abstract: Automated question answering (QA) over electronic health records (EHRs) demands precise evidence retrieval, faithful answer generation, and explicit gro
arXiv:2605.09301v1 Announce Type: cross Abstract: The Capacitated Vehicle Routing Problem (CVRP) underpins modern last-mile logistics. Current Neural Combinatorial Optimization (NCO) methods construct
arXiv:2605.05373v2 Announce Type: replace Abstract: A key capability of intelligent agents is operating under partial observability: reasoning and acting effectively despite missing or incomplete stat
arXiv:2605.09939v1 Announce Type: new Abstract: Autonomous and safe navigation of tractor-trailer systems requires accurate, real-time collision avoidance and dynamically feasible control, particularl
arXiv:2605.09316v1 Announce Type: cross Abstract: Query-separated computation forces a representation to play an operational role: data are encoded before a query is known, and a later decoder can ans
arXiv:2605.08179v1 Announce Type: cross Abstract: Radar sounders are electromagnetic instruments that can probe deep into the subsurface of Earth and other planetary bodies by processing the echo of t
arXiv:2506.01250v2 Announce Type: replace Abstract: In this paper, we address the contextual dueling bandit problem by proposing variance-aware algorithms that leverage neural networks to approximate
arXiv:2605.10878v1 Announce Type: new Abstract: Why does weight decay work? We prove that, in any fixed-precision regime, the smallest weight norm of a looped neural network outputting a binary string
arXiv:2605.08495v1 Announce Type: new Abstract: Deep learning and large public datasets have recently catalyzed the proliferation of AI models for processing brain recordings. However, systematically
arXiv:2605.08458v1 Announce Type: new Abstract: Coherent, continuous spatial representations are critical for synthesizing physical and perceptual phenomena into a single representational space. Radia
arXiv:2605.08192v1 Announce Type: cross Abstract: Frontier AI safety claims - published assertions that a highly capable general-purpose model is below a threshold of concern, adequately mitigated, or
arXiv:2605.08373v1 Announce Type: cross Abstract: Recent advances in neuroimaging have deepened our understanding of the brain's complex functional and structural organization. Among these, functional
arXiv:2605.10675v1 Announce Type: new Abstract: Event cameras offer distinct advantages over conventional frame-based sensors, including microsecond-level temporal resolution, high dynamic range, and
arXiv:2605.09595v1 Announce Type: cross Abstract: Reinforcement learning (RL) has enabled robust quadruped locomotion over complex terrain, but most learned controllers are trained offline with backpr
arXiv:2509.21671v2 Announce Type: replace Abstract: High-resolution neural datasets enable foundation models for the next generation of brain-computer interfaces and neurological treatments. The commu
arXiv:2605.08188v1 Announce Type: cross Abstract: Human attention is the gateway to conscious perception, memory and decision-making. However, its role in modern transformer models remains largely une
arXiv:2605.10804v1 Announce Type: new Abstract: Campus well-being underpins academic success, yet many universities lack effective methods for monitoring satisfaction and detecting mental health risks
arXiv:2505.20001v5 Announce Type: replace Abstract: Multi-modal object Re-IDentification (ReID) aims to obtain complete identity features across heterogeneous modalities. However, most existing method
arXiv:2605.09387v1 Announce Type: new Abstract: While Large Language Models (LLMs) have catalyzed progress in embodied intelligence, a fundamental gap between their inherent probabilistic uncertainty
arXiv:2605.08452v1 Announce Type: new Abstract: The ability to derive precise spatial and physical insights is a cornerstone of vision-language models (VLMs), yet their poor performances in related sp
arXiv:2602.04549v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) revolutionized novel view rendering. Instead of inferring from dense spatial points, as implicit representations do, 3D
arXiv:2510.20797v2 Announce Type: replace-cross Abstract: Context compression reduces Transformer inference costs by replacing lengthy inputs with shorter pre-computed representations. It carries sign
arXiv:2605.09328v1 Announce Type: new Abstract: Pre-trained text-to-image (T2I) diffusion models have shown strong potential for real-world image super-resolution (Real-ISR), owing to their noise-star
arXiv:2605.08144v1 Announce Type: cross Abstract: Diffusion models have achieved remarkable success across a wide range of generative tasks, yet their training paradigm largely treats injected noise a
arXiv:2605.08221v1 Announce Type: cross Abstract: This paper presents NoisyCoconut, a novel inference-time method that enhances large language model (LLM) reliability by manipulating internal represen
arXiv:2605.08306v1 Announce Type: cross Abstract: Body composition assessment (BCA) provides detailed information about the distribution of different tissue types in the body, enabling more precise ch
arXiv:2605.08913v1 Announce Type: cross Abstract: Autoregressive inference is typically assumed to scale predictably with decoding length, and key-value (KV) caching is widely regarded as a universall
arXiv:2605.08999v1 Announce Type: new Abstract: In machine learning, a critical class of decision-related problems concerns preventing predicted undesirable outcomes, referred to as the extit{avoiding
arXiv:2605.09058v1 Announce Type: cross Abstract: We introduce Nonlinear GENERIC Informed Neural Networks (N-GINNs), a deep learning framework for discovering evolution equations of systems governed b
arXiv:2605.10823v1 Announce Type: new Abstract: Reversible instance normalization (RevIN) and its successors (Dish-TS, SAN, FAN) have become the de facto plug-in for time-series forecasting, yet the m
arXiv:2605.08193v1 Announce Type: cross Abstract: Normalization Equivariance (NE), equivariance to global contrast and brightness transforms, improves robustness to distribution shift in image-to-imag
arXiv:2605.10379v1 Announce Type: new Abstract: Large language models (LLMs) have become capable mathematical problem-solvers, often producing correct proofs for challenging problems. However, correct
arXiv:2605.09490v1 Announce Type: new Abstract: Reasoning LLMs produce thousands of chain-of-thought tokens whose KV cache must reside in scarce GPU HBM. The dominant response -- permanently evicting
arXiv:2605.08778v1 Announce Type: new Abstract: Deploying LLMs in multi-turn dialogues facilitates jailbreak attacks that distribute harmful intent across seemingly benign turns. Recent training-based
arXiv:2605.10676v1 Announce Type: new Abstract: During MLLM decoding, attention often abnormally concentrates on irrelevant image tokens. While existing research dismisses this as invalid noise and fo
arXiv:2605.10061v1 Announce Type: cross Abstract: Futrell and Mahowald (2025) frame the success of neural language models (LMs) as supporting gradient, usage-based linguistic theories. I argue that LM
arXiv:2510.03895v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models represent a pivotal advance in embodied intelligence, yet they confront critical barriers to real-world de
arXiv:2605.09950v1 Announce Type: cross Abstract: Most feature selection algorithms, especially wrapper methods, run inefficiently on CPU based platforms because of their high computational complexity