VASAE: Naming SAE Dictionary Directions with Vocabulary-Aligned Anchoring
DGX agentarXiv:2606.27941v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) provide useful decompositions of Transformer residual streams, but their learned features are usually named post hoc rather