Evaluating and Calibrating LLM Confidence on Questions with Multiple Correct Answers
DGX agentarXiv:2602.07842v2 Announce Type: replace Abstract: Confidence calibration is essential for making large language models (LLMs) reliable, yet existing training-free methods have been primarily studied