Rethinking Uncertainty Evaluation in Large Language Models
DGX agentarXiv:2607.19367v1 Announce Type: new Abstract: Calibration is the primary criterion for evaluating LLM confidence, but it is insufficient: it admits trivially incoherent estimators, depends on the ev