Look But Don't Touch with Sparse Autoencoders for Unlearning in Diffusion Models
DGX agentarXiv:2606.31699v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) have recently been proposed as interpretable tools for concept-level manipulation, under the assumption that isolated featu