Can an MLP Absorb Its Own Skip Connection?
arXiv:2604.23705v1 Announce Type: new Abstract: We study when a skip connection around a single-hidden-layer MLP can be absorbed into a residual-free MLP of the same width. We first show that for any
Knowledge catalogue
arXiv:2604.23705v1 Announce Type: new Abstract: We study when a skip connection around a single-hidden-layer MLP can be absorbed into a residual-free MLP of the same width. We first show that for any
arXiv:2604.23471v1 Announce Type: cross Abstract: This study asks whether the threat of AI detection changes how people write with AI, and whether other people can tell the difference. In a two-phase
arXiv:2604.24444v1 Announce Type: new Abstract: Despite the growing use of large language models (LLMs) for writing tasks, users may hesitate to rely on LLMs when personal style is important. Post-edi
arXiv:2601.06352v2 Announce Type: replace Abstract: Adapting large language models to individual users remains challenging due to the tension between fine-grained personalization and scalable deployme
arXiv:2604.23876v1 Announce Type: new Abstract: We present Cardiac Stability Theory (CST), an axiomatically grounded framework formally defining cardiovascular health as a stability margin around a ca
arXiv:2604.23718v1 Announce Type: new Abstract: As dental caries appear as subtle, low-contrast lesions in intraoral imaging, existing deep learning models face significant challenges in the early det
arXiv:2604.23800v1 Announce Type: new Abstract: Causal representation learning aims to recover the latent causal variables and their causal relations, typically represented by directed acyclic graphs
arXiv:2604.23628v1 Announce Type: cross Abstract: Hierarchical clustering is a fundamental task in data analysis, yet for a long time it lacked a principled objective function. Dasgupta [STOC 2016] in
arXiv:2604.24447v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are promising for generalist robot control, but on-robot deployment is bottlenecked by real-time inference under t
arXiv:2510.13312v2 Announce Type: replace Abstract: We present ChatR1, a reasoning framework based on reinforcement learning (RL) for conversational question answering (CQA). Reasoning plays an import
arXiv:2604.22989v1 Announce Type: cross Abstract: Recent medical multimodal foundation models are built as multimodal LLMs (MLLMs) by connecting a CLIP-pretrained vision encoder to an LLM using LLaVA-
arXiv:2604.05500v2 Announce Type: replace Abstract: Nighttime image dehazing faces a more complex degradation pattern than its daytime counterpart, as haze scattering couples with low illumination, no
arXiv:2601.22301v2 Announce Type: replace Abstract: Traditional rendering pipelines rely on complex assets, accurate materials and lighting, and substantial computational resources to produce realisti
arXiv:2509.11717v5 Announce Type: replace-cross Abstract: Text-guided sound separation enables flexible audio editing, assistive listening, and open-domain source extraction, but systems such as Audio
arXiv:2604.24634v1 Announce Type: cross Abstract: Light-activated drugs are a promising way to treat localized diseases for which existing treatments have severe side effects. However, their developme
arXiv:2604.23952v1 Announce Type: cross Abstract: Stochastic reduced-order models are widely used to represent the effective dynamics of complex systems, but estimating their drift and diffusion coeff
arXiv:2604.22787v1 Announce Type: cross Abstract: Africa's green industrialization imperative demands reliable infrastructure for monitoring air quality. We present a satellite-reanalysis PM2.5 fusion
arXiv:2309.00578v2 Announce Type: replace Abstract: In the context of unsupervised learning, Lloyd's algorithm is one of the most widely used clustering algorithms. It has inspired a plethora of work
arXiv:2604.22801v1 Announce Type: cross Abstract: It is a challenging task to forecast equity prices in fast moving financial markets as this becomes even more difficult when the predictive signal is
arXiv:2604.24693v1 Announce Type: new Abstract: Linear activation steering is a powerful approach for eliciting the capabilities of large language models and specializing their behavior using limited
arXiv:2604.23987v1 Announce Type: new Abstract: Continual learning for large language models is typically evaluated through accuracy retention under sequential fine-tuning. We argue that this perspect
arXiv:2410.13903v3 Announce Type: replace-cross Abstract: Proprietary large language models (LLMs) exhibit strong generalization capabilities across diverse tasks and are increasingly deployed on edge
arXiv:2604.24170v1 Announce Type: new Abstract: Concept Bottleneck Models (CBMs) predict through human-interpretable concepts, but they typically output point concept probabilities that conflate epist
arXiv:2604.24361v1 Announce Type: new Abstract: Large language models (LLMs) have achieved strong performance in general machine translation, yet their ability in culture-aware scenarios remains poorl
arXiv:2604.23693v1 Announce Type: new Abstract: Heterogeneous multi-robot systems feature significant adaptability for complex environments. However, effective collaboration that fully exploits the ro
arXiv:2504.09499v2 Announce Type: replace-cross Abstract: Hattrick is a free web-based probabilistic football manager game with over 200,000 users competing for titles at national and international le
arXiv:2603.28463v2 Announce Type: replace Abstract: Domain generalization in fundus imaging is challenging due to variations in acquisition conditions across devices and clinical settings. The inabili
arXiv:2604.22909v1 Announce Type: new Abstract: Understanding and representing complex climate variability is essential for both scientific analysis and predictive modeling. However, identifying meani
arXiv:2604.24647v1 Announce Type: cross Abstract: Long-context reasoning is a critical capability of large language models (LLMs), enabling applications such as long-document understanding, summarizat
arXiv:2604.23976v1 Announce Type: new Abstract: The sense of family connectedness may support positive outcomes including individual well-being, resilience, and healthy family functioning. However, as
arXiv:2601.02455v2 Announce Type: replace-cross Abstract: Deploying Automatic Speech Recognition (ASR) models on memory-constrained edge devices requires aggressive low-bit weight quantization. Layer-
arXiv:2604.24547v1 Announce Type: new Abstract: Progression to dialysis or end-stage renal disease is a rare but clinically important outcome. Clinicians need evidence on how medication exposures infl
arXiv:2604.24719v1 Announce Type: new Abstract: Segmentation models such as Segment Anything Model (SAM) and SAM2 achieve strong prompt-driven zero-shot performance. However, their training on natural
arXiv:2604.24692v1 Announce Type: new Abstract: We propose Noise-Based Spectral Embedding (NBSE), a physics-informed framework for selecting informative features from high-dimensional data without gre
arXiv:2604.24575v1 Announce Type: new Abstract: Diffusion models are primarily trained for image synthesis, yet their denoising trajectories encode rich, spatially aligned visual priors. In this paper
arXiv:2604.23636v1 Announce Type: new Abstract: In this work, we study Source-Free Unsupervised Domain Adaptation under corruption-induced domain shifts, where performance degradation is caused by nat
arXiv:2511.21715v2 Announce Type: replace-cross Abstract: This paper argues that dataset structure is important in image recognition tasks (among other tasks). Specifically, we focus on the nature and
arXiv:2604.22847v1 Announce Type: new Abstract: We introduce Dream-Cubed, a large-scale dataset of Minecraft worlds at voxel resolution, and a family of models using cubes as powerful compositional un
arXiv:2505.13975v4 Announce Type: replace Abstract: While Large Reasoning Models (LRMs) have demonstrated success in complex reasoning tasks through long chain-of-thought (CoT) reasoning, their infere
arXiv:2604.24663v1 Announce Type: cross Abstract: We study finite-horizon quadratic control of linear systems with bilinear observations, in which the control input affects not only the state dynamics
arXiv:2603.10926v1 Announce Type: cross Abstract: Time-series anomaly detectors are commonly compared on workstation-class hardware under unconstrained execution. In-vehicle monitoring, however, requi
arXiv:2604.23763v1 Announce Type: new Abstract: Large diffusion transformers (DiTs) follow global editing instructions well but consistently leak local edits into unrelated regions, because joint-atte
arXiv:2604.22992v1 Announce Type: new Abstract: Reliable object perception is necessary for general-purpose service robots. Open-vocabulary detectors struggle to generalize beyond a few classes and fu
arXiv:2604.24555v1 Announce Type: new Abstract: We consider online learning problems under a partial observability model capturing situations where the information conveyed to the learner is between f
arXiv:2604.23172v1 Announce Type: new Abstract: In this work, we developed and tested 3 techniques for vector quantization (VQ) based model weight compression. To mitigate codebook collapse and enable
arXiv:2604.23532v1 Announce Type: cross Abstract: Short-term human pose prediction plays a crucial role in interactive systems, assistive robots, and emotion-aware human-computer interaction[1-3]. Whi
arXiv:2601.00823v2 Announce Type: replace Abstract: Large reasoning models (LRMs) have heterogeneous inference energy costs based on which model is used and how much it reasons. To reduce energy, it i
arXiv:2604.23426v1 Announce Type: new Abstract: Federated learning (FL) is a distributed machine learning method where multiple devices collaboratively train a model under the management of a central
arXiv:2604.24563v1 Announce Type: cross Abstract: Machine-learning interatomic potentials (MLIPs) have enabled molecular dynamics at near ab initio accuracy, yet remain limited to energies and forces
arXiv:2604.22776v1 Announce Type: cross Abstract: A chef's intuition about flavor, texture, and cultural identity represents tacit knowledge that is difficult to articulate yet central to culinary pra
arXiv:2604.22855v1 Announce Type: new Abstract: The core objective of image captioning is to achieve lossless semantic compression from visual signals into textual modalities. However, the reliance on
arXiv:2510.02629v3 Announce Type: replace Abstract: Context utilisation, the ability of Language Models (LMs) to incorporate relevant information from the provided context when generating responses, r
arXiv:2604.24609v1 Announce Type: new Abstract: Many sign language translation (SLT) systems operate on pose sequences instead of raw video to reduce input dimensionality, improve portability, and par
arXiv:2604.23887v1 Announce Type: cross Abstract: LLM-powered applications routinely embed secrets in system prompts, yet models can be tricked into revealing them. We built an adaptive attacker that
arXiv:2505.17855v2 Announce Type: replace Abstract: Understanding sources of a model's uncertainty regarding its predictions is crucial for effective human-AI collaboration. Prior work proposes using
arXiv:2604.23260v1 Announce Type: cross Abstract: An approach to construct explicit integral representations for two-layer ReLU networks is presented, which provides relatively simple representations
arXiv:2604.24706v1 Announce Type: cross Abstract: Learning-based control techniques use data from past trajectories to control systems with uncertain dynamics. However, learning-based controllers are
arXiv:2604.23860v1 Announce Type: cross Abstract: Egocentric videos provide a distinctive setting in which sound serves as crucial cues to understand user activities and surroundings, particularly whe
arXiv:2501.02673v4 Announce Type: replace Abstract: Having a sufficient quantity of quality data is a critical enabler of training effective machine learning models. Being able to effectively determin
arXiv:2604.23562v1 Announce Type: cross Abstract: The relationship between brain lateralization and cognitive functions is well-documented. The left hemisphere primarily handles tasks such as language