Subspace-Aware Sparse Autoencoders for Effective Mechanistic Interpretability
DGX agentarXiv:2606.06333v1 Announce Type: cross Abstract: Sparse Autoencoders (SAEs) are widely used for mechanistic interpretability in large language models, yet their formulation assigns each latent featur