ConfidenceBench: Evaluating Confidence Calibration in Large Language Models
DGX agentarXiv:2607.20526v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in settings where fluent but incorrect answers can be costly. In these settings, accuracy alone i