Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization
arXiv:2605.21801v1 Announce Type: cross Abstract: Post-training has become central to improving reasoning and alignment in large language models, where critic-free models enable scalable learning from