How Optimality Structures Sparse Dictionaries: A Theory for Understanding SAE Representations
arXiv:2606.02385v1 Announce Type: cross Abstract: Sparse Autoencoders (SAEs) have found success parsing neural representations into interpretable concepts, providing a basis for understanding and cont