ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment
arXiv:2601.21484v2 Announce Type: replace Abstract: Reinforcement Learning (RL) post-training alignment for language models is effective, but also costly and unstable in practice, owing to its complic