When Structured Sparse Autoencoders Learn Consistent Concepts Across Modalities
DGX agentarXiv:2607.08605v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) have emerged as a promising technique for mechanistic interpretability by learning a set of sparse latent features in large