ReSAE: Residualized Sparse Autoencoders for Multi-Layer Transformer Interventions
DGX agentarXiv:2605.27819v1 Announce Type: cross Abstract: Sparse autoencoders are usually trained one layer at a time, even though transformer residual stream activations are strongly coupled across depth. Th