Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure
DGX agentarXiv:2605.08740v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) decompose transformer residual streams into interpretable feature dictionaries, yet the relationship between SAE width and