Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
arXiv:2605.13537v1 Announce Type: cross Abstract: Inference-time alignment techniques offer a lightweight alternative or complement to costly reinforcement learning, while enabling continual adaptatio