Safety
Explaining Intrinsic Moral Self-Correction with Mechanistic Interpretability
arXiv:2505.11924v4 Announce Type: replace-cross Abstract: Intrinsic moral self-correction refers to the phenomenon where a language model refines its ethical judgments or aligns its outputs purely thr
arXiv:2505.11924v4 Announce Type: replace-cross Abstract: Intrinsic moral self-correction refers to the phenomenon where a language model refines its ethical judgments or aligns its outputs purely through prompting. While effective across diverse tasks, its mechanism remains unclear. We hypothesize intrinsic moral self-correction functions by steering hidden representations along interpretable latent directions. Evaluating six LLMs across four morality-related tasks, we demonstrate that the representation shifts induced by self-correction prompts align with contrastive steering vectors. This alignment transfers even when the steering vectors are constructed from a disjoint corpus. Notably, when applied via activation addition, these prompt-induced shifts can alter model behavior more effectively than the self-correction prompts and the steering vectors. Our findings suggest representation steering is the mechanistic driver of intrinsic moral self-correction.
Related
- Dual Mechanisms of Value Expression: Intrinsic vs. Prompted Values in Large Language Models
- Evolutionary Guided Decoding: Iterative Value Refinement for LLMs
- Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs
- Framing Instability in LLM Ethical Stance: Auditing Negation Sensitivity in Moral Dilemmas
Source: arXiv cs.AI | 2026-08-24