Tutorials
Revenge of Monosemanticity: Specialized Neurons Improve Data Efficiency in MLPs
arXiv:2608.24007v1 Announce Type: new Abstract: Understanding how neural networks learn and organize features is central to understanding their behavior. Much existing theory of feature learning has f
arXiv:2608.24007v1 Announce Type: new Abstract: Understanding how neural networks learn and organize features is central to understanding their behavior. Much existing theory of feature learning has focused on the emergence of a global low-dimensional predictive geometry. We show that this picture is incomplete. In regression problems with clustered data, we demonstrate that multilayer perceptrons (MLPs) naturally develop monosemantic specialized neurons: individual neurons become strongly aligned with a specific predictive feature relevant to a particular region of the input space. Rather than learning a single global low-dimensional representation, MLPs learn a collection of local low-dimensional representations that can collectively span a high-dimensional space. This specialization provably gives MLPs a data-efficiency advantage over feature-learning methods based on a global low-dimensional representation.
Related
- Deep Learning as Neural Low-Degree Filtering: A Spectral Theory of Hierarchical Feature Learning
- Shallow Neural Networks Learn Low-Degree Spherical Polynomials with Feature Learning by Learnable Channel Attention
- Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders
- Balancing Learning Rates Across Layers: Exact Two-Step Dynamics and Optimal Scaling in Linear Neural Networks
Source: arXiv cs.LG | 2026-08-26