UCPO: Uncertainty-Aware Policy Optimization
DGX agentarXiv:2601.22648v2 Announce Type: replace Abstract: The key to building trustworthy large language models (LLMs) lies in endowing them with inherent uncertainty expression capabilities, thereby mitiga