Muown: Row-Norm Control for Muon Optimization
arXiv:2605.10797v1 Announce Type: new Abstract: Muon has emerged as a strong competitor to AdamW for language model pre-training, yet its behavior at scale is sensitive to weight decay. Recent work ha
Knowledge catalogue
arXiv:2605.10797v1 Announce Type: new Abstract: Muon has emerged as a strong competitor to AdamW for language model pre-training, yet its behavior at scale is sensitive to weight decay. Recent work ha
arXiv:2507.14958v2 Announce Type: replace Abstract: Current models have achieved impressive performance on reasoning-intensive tasks, yet optimizing their reasoning efficiency remains an open challeng
arXiv:2605.09863v1 Announce Type: cross Abstract: Production LLM coding agents drift over long sessions: they forget user-specified constraints, slip into mistakes the user already flagged, and confab
arXiv:2605.09176v1 Announce Type: cross Abstract: Training large language models requires optimization algorithms that are not only statistically effective, but also computationally and memory efficie
arXiv:2605.08179v1 Announce Type: cross Abstract: Radar sounders are electromagnetic instruments that can probe deep into the subsurface of Earth and other planetary bodies by processing the echo of t
arXiv:2605.08188v1 Announce Type: cross Abstract: Human attention is the gateway to conscious perception, memory and decision-making. However, its role in modern transformer models remains largely une
arXiv:2605.08452v1 Announce Type: new Abstract: The ability to derive precise spatial and physical insights is a cornerstone of vision-language models (VLMs), yet their poor performances in related sp
arXiv:2605.09055v1 Announce Type: cross Abstract: Recent agentic-robotics systems, from Code-asPolicies to modern vision-language-action (VLA) foundation models, presuppose that drivers, SDKs, or ROS-
arXiv:2605.09433v1 Announce Type: new Abstract: Existing preference datasets for text-to-image models typically store only the final winner/loser images. This representation is insufficient for rectif
arXiv:2605.09075v1 Announce Type: cross Abstract: Although the Laplace approximation offers a simple route to uncertainty quantification in deep neural networks, its reliance on inverting large Hessia
arXiv:2605.08660v1 Announce Type: new Abstract: In the recent literature, Support Vector Regression (SVR) has been cited as one of the weakest performers on the California Housing benchmark dataset, w
arXiv:2502.06830v5 Announce Type: replace-cross Abstract: Probabilistic intraday electricity price forecasting is becoming increasingly important for short-term power-system operation. With increasing
arXiv:2605.09036v1 Announce Type: new Abstract: Accurate and efficient storm-surge emulation is essential for coastal hazard assessment, yet high-fidelity hydrodynamic models remain too expensive for
arXiv:2602.06733v2 Announce Type: replace-cross Abstract: Multi-Agent Path Finding (MAPF) is a representative multi-agent coordination problem, where multiple agents are required to navigate to their
arXiv:2605.10219v1 Announce Type: cross Abstract: We study the parameterized complexity of testing approximate first-order stationarity at a prescribed point for continuous piecewise-affine (PA) funct
arXiv:2605.09636v1 Announce Type: new Abstract: PDE-to-solver code generation aims to automatically synthesize executable numerical solvers from partial differential equation (PDE) specifications. Thi
arXiv:2605.09339v1 Announce Type: cross Abstract: Human color categories are not uniformly distributed in perceptual space, yet most computational color models still assume fixed and evenly structured
arXiv:2605.09638v1 Announce Type: new Abstract: Ensuring the security of reinforcement learning (RL) models is critical, particularly when they are trained by third parties and deployed in real-world
arXiv:2605.10275v1 Announce Type: new Abstract: Polarimetric imaging captures surface polarization characteristics, such as the Degree of Linear Polarization (DoLP) and the Angle of Polarization (AoP)
arXiv:2605.10581v1 Announce Type: new Abstract: Retinal vessel segmentation is crucial for diagnosis and assessment of ocular diseases. Notably, segmentation of small retinal vessels has been consiste
arXiv:2605.09365v1 Announce Type: new Abstract: Enterprise workloads are dominated by deterministic, structured, and knowledge-dependent tasks operating under strict cost, latency, and reliability con
arXiv:2605.08687v1 Announce Type: cross Abstract: Data preparation is a central and time-consuming stage in data analysis workflows. Traditionally, commercial tools have relied on graphical user inter
arXiv:2605.09749v1 Announce Type: new Abstract: Discrete diffusion models generate structured sequences by progressively unmasking tokens, but enforcing global property constraints during generation r
arXiv:2605.10614v1 Announce Type: new Abstract: Multi-agent LLM systems introduce a security risk in which sensitive information accessed by one agent can propagate through shared context and reappear
arXiv:2510.25372v2 Announce Type: replace Abstract: Visual Prompt Tuning (VPT) of pre-trained Vision Transformers (ViTs) has proven highly effective as a parameter-efficient fine-tuning technique for
arXiv:2605.09431v1 Announce Type: new Abstract: Cryptocurrency pump-and-dump schemes coordinated via Telegram threaten market integrity. However, existing research addressing this specific threat has
arXiv:2605.09808v1 Announce Type: new Abstract: User simulators are increasingly leveraged to build interactive AI assistants, yet how to measure the quality of these simulators remains an open questi
arXiv:2605.08515v1 Announce Type: new Abstract: Unlike standard expected-return Reinforcement Learning (RL), Distributional RL (DRL) models the full return distribution, making it better-suited for un
arXiv:2605.08170v1 Announce Type: new Abstract: Neural operators have emerged as a powerful tool for learning mappings between infinite-dimensional function spaces. However, their approximation proper
arXiv:2605.08423v1 Announce Type: cross Abstract: We present a data-adaptive method for parameter-efficient fine-tuning of large neural networks. Standard low-rank adaptation methods improve efficienc
arXiv:2603.11566v2 Announce Type: replace Abstract: 4D radar-camera sensing configuration has gained increasing importance in autonomous driving. However, existing 3D object detection methods that fus
arXiv:2605.08857v1 Announce Type: new Abstract: Recent advances in uncertainty quantification for time series forecasting show that conformal prediction can provide reliable prediction intervals, yet
arXiv:2605.08454v1 Announce Type: cross Abstract: Recovering continuous-time dynamics from discrete observations is difficult because local supervision (e.g., pointwise regression targets, derivative
arXiv:2605.10317v1 Announce Type: cross Abstract: Knowledge graph embedding (KGE) models typically represent each relation as an operator on entity embeddings. In this work, we identify three structur
arXiv:2605.08840v1 Announce Type: new Abstract: Large language models (LLMs) face growing challenges in efficient generative inference due to the increasing memory demands of Key-Value (KV) caches, es
arXiv:2605.09905v1 Announce Type: cross Abstract: Automatic sleep staging commonly adopts Transformers under the assumption that they learn complex long-range dependencies. We challenge this view by r
arXiv:2511.21600v2 Announce Type: replace-cross Abstract: The rise of generative AI has enabled the production of high-fidelity synthetic tabular data across fields such as healthcare, finance, and pu
arXiv:2605.10357v1 Announce Type: cross Abstract: Multimodal misinformation increasingly leverages visual persuasion, where repurposed or manipulated images strengthen misleading text. We introduce ex
arXiv:2605.08391v1 Announce Type: new Abstract: Cooperative multi-agent reinforcement learning agents that act on partial local observations face a fundamental information bottleneck: the knowledge ne
In today's hyper-connected market, an enterprise's most valuable asset — mission-critical data — often remains trapped in legacy silos. For years, leadership teams have navigated a data pipeline dilem
arXiv:2605.09038v1 Announce Type: new Abstract: Teaching language models to use search tools is not only a question of whether they search, but also of whether they issue good queries. This is especia
arXiv:2512.19219v2 Announce Type: replace-cross Abstract: Low-rank adaptation (LoRA) is widely used for parameter-efficient fine-tuning, but its standard all-token, all-head design ignores the heterog
arXiv:2605.08874v1 Announce Type: new Abstract: Open-vocabulary semantic segmentation requires adapting image-level vision-language models such as CLIP to dense pixel-level prediction, which is challe
arXiv:2605.09271v1 Announce Type: new Abstract: Although natural language is the default medium for Large Language Models (LLMs), its limited expressive capacity creates a profound bottleneck for comp
arXiv:2605.09070v1 Announce Type: cross Abstract: Many jailbreak attack research papers report attack success rates for a limited number of parameter settings, even though there are many combinations
arXiv:2605.08658v1 Announce Type: cross Abstract: SKETCHVERIFY is a within-tier cost-performance policy, not a universal accuracy improvement. The operational question: a practitioner stuck with a sma
Mike Isaac / New York Times: Sources: Anthropic is in talks to raise between 30B and 50B in a funding round that would value it at up to 950B — The start-up, which recently released a powerful A.I. mo
arXiv:2605.09403v1 Announce Type: cross Abstract: Architectural choices inside the Transformer feedforward network (FFN) block do not merely affect the block itself; they reshape the computations lear
arXiv:2605.08220v1 Announce Type: new Abstract: The automated extraction of data from scientific charts is a critical task for large-scale literature analysis. While multimodal Large Language Models (
arXiv:2603.00541v2 Announce Type: replace Abstract: Generative foundation models are increasingly scaled in both width and depth, posing significant challenges for stable feature learning and reliable
arXiv:2605.10674v1 Announce Type: cross Abstract: Rejection Fine-Tuning (RFT) is a standard method for training LLM agents, where unsuccessful trajectories are discarded from the training set. In the
arXiv:2605.09415v1 Announce Type: new Abstract: The growing integration of AI into cybersecurity is reshaping the balance between attackers and defenders. When access to advanced AI-enabled defence to
arXiv:2506.01352v2 Announce Type: replace Abstract: Decentralized training of large language models offers the opportunity to pool computational resources across geographically distributed participant
arXiv:2505.11604v5 Announce Type: replace Abstract: Editing presentation slides is a frequent yet tedious task, ranging from creative layout design to repetitive text maintenance. While recent GUI-bas
arXiv:2508.05463v2 Announce Type: replace-cross Abstract: Neural networks excel across a wide range of tasks, yet remain black boxes. In particular, how their internal representations are shaped by th
arXiv:2605.05775v2 Announce Type: replace-cross Abstract: We report the design and results of the third autoPET challenge (MICCAI 2024), which benchmarked automated lesion segmentation in whole-body P
arXiv:2605.10698v1 Announce Type: cross Abstract: Multi-agent systems (MAS) assume that collaborating inherently improves Large Language Model (LLM) reasoning. We challenge this by demonstrating that
arXiv:2605.10799v1 Announce Type: cross Abstract: Corruption studies, the primary tool for evaluating chain-of-thought (CoT) faithfulness, identify which chain positions are 'computationally important
The silent removal of Study Mode from ChatGPT is a big mistake (both Claude and Gemini still have theirs) We have enough evidence that using AI in assistant mode to study can hurt learning because it
arXiv:2605.10117v1 Announce Type: cross Abstract: Autonomous driving scenes range from empty highways to dense intersections with dozens of interacting road users, yet current 3D detection models appl