Graph-Regularized Sparse Autoencoders for LLM Safety Steering
DGX agentarXiv:2512.06655v3 Announce Type: replace-cross Abstract: Sparse autoencoders (SAEs) are increasingly used to extract activation directions for inference-time steering, but their standard sparsity obj