Alignment Makes Language Models Normative, Not Descriptive
DGX agentarXiv:2603.17218v2 Announce Type: replace-cross Abstract: Post-training alignment optimizes language models to match human preference signals, but this objective is not equivalent to modeling observed