Safety
Running the same CNA search on base models (before instruction tuning) yields a structurally similar set of neurons, but ablating them produ…
Running the same CNA search on base models (before instruction tuning) yields a structurally similar set of neurons, but ablating them produces almost no behavioral change. We read this as evidence th
Running the same CNA search on base models (before instruction tuning) yields a structurally similar set of neurons, but ablating them produces almost no behavioral change. We read this as evidence that the refusal mechanism is not latent in the pretrained model. The structural substrate is there, but alignment fine-tuning is what wires it up as a behavioral gate.
Source: Nous Research (X) | 2026-05-19