Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection
DGX agentarXiv:2605.23036v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) enable feature-level mechanistic interpretability and activation steering in large language models (LLMs), but SAE-based lang