Safe RLHF Beyond Expectation: Stochastic Dominance for Universal Spectral Risk Control
DGX agentarXiv:2603.10938v2 Announce Type: replace-cross Abstract: Safe Reinforcement Learning from Human Feedback (RLHF) typically enforces safety through expected cost constraints, but the expectation captur