ICA Lens: Interpreting Language Models Without Training Another Dictionary
DGX agentarXiv:2606.11722v1 Announce Type: cross Abstract: Finding interpretable directions in language-model representations is critical for understanding and controlling model behavior. Sparse autoencoders (