Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning
arXiv:2608.00301v1 Announce Type: cross Abstract: Error-penalized scoring rules (+1 for a correct answer, -lambda for a wrong one, 0 for abstaining) are increasingly prescribed against hallucination: