Finding Usable Weight Mechanisms with Tiled SVD
DGX agentarXiv:2608.06969v1 Announce Type: new Abstract: The dominant approach to mechanistic interpretability trains proxy dictionaries such as sparse autoencoders and labels features from max-activating text