Safety
What we find most useful about CNA is that the intervention is simple yet powerful. The steering is a multiplicative ablation on a sparse se…
What we find most useful about CNA is that the intervention is simple yet powerful. The steering is a multiplicative ablation on a sparse set of MLP neurons, which makes CNA a clean addition on top of
What we find most useful about CNA is that the intervention is simple yet powerful. The steering is a multiplicative ablation on a sparse set of MLP neurons, which makes CNA a clean addition on top of standard production pipelines: instruction tuning, RL, and safety post-training. We find the neuron basis to be ripe for further exploration in interpretability and steering domains. Paper: https://arxiv.org/abs/2605.12290 Blog: https://nousresearch.com/neuron-steering Code: https://github.com/NousResearch/neural-steering HF: https://huggingface.co/papers/2605.12290 CNA is a contribution from the mechanistic interpretability team at Nous. If you want to work on problems like this, find us on Discord.
Source: Nous Research (X) | 2026-05-19