Beyond Linear Activation Steering: Invertible Latent Transformations for Controlling LLM Behavior
DGX agentarXiv:2606.08454v1 Announce Type: new Abstract: Activation steering provides a lightweight inference-time mechanism for controlling large language models (LLMs) by modifying their internal activation