Research
Different Facets of Verbalised Overconfidence: an Interpretability Study
arXiv:2608.18106v1 Announce Type: cross Abstract: Large language models tend to overconfidence, giving assertive answers when the evidence suggests hedging or abstention. Using controlled reasoning sc
arXiv:2608.18106v1 Announce Type: cross Abstract: Large language models tend to overconfidence, giving assertive answers when the evidence suggests hedging or abstention. Using controlled reasoning scenarios that manipulate logical necessity and possibility, we study this behavior in Qwen3-4B, across three ways to express uncertainty: verbal epistemic markers, abstention, and numeric confidence scores. Our results confirm this tendency toward overconfidence, particularly when the model is prompted to output a numeric confidence score. At the interpretability level, we propose a method that differentially identifies transcoder features responsible for uncertainty and certainty. Our analysis reveals Qwen3-4B's default mechanism favors certainty generation through a broad coalition of shared features, while uncertainty is implemented as a sparse override mediated by a small set of dedicated features. Intervening on these uncertainty features both causally proves this imbalance underlying overconfidence and also mitigate overconfident errors. The same set of features generalise across the three uncertainty-expression settings, languages, and an out-of-distribution modality task.
Related
- When Confidence Fails: Overconfidence in LLMs under Uncertainty and Missing Clinical Information
- I-CALM: Incentivizing Confidence-Aware Abstention for LLM Selective Answering
- Stable Miscalibration in Large Language Models: A Practical View of High-Confidence Errors
Source: arXiv cs.AI | 2026-08-20