Model Releases
Concurrent Criterion Validation of a Validity Screen for LLM Confidence Signals via Selective Prediction
arXiv:2604.17716v1 Announce Type: new Abstract: The validity screen (Cacioli, 2026d, 2026e) classifies LLM confidence signals as Valid, Indeterminate, or Invalid. We test whether these classifications
arXiv:2604.17716v1 Announce Type: new Abstract: The validity screen (Cacioli, 2026d, 2026e) classifies LLM confidence signals as Valid, Indeterminate, or Invalid. We test whether these classifications predict selective prediction performance. Twenty frontier LLMs from seven families were evaluated on 524 items across six cognitive tracks. Valid models show mean Type 2 AUROC = .624 (SD = .048). Invalid models show mean AUROC = .357 (SD = .231). Cohen's d = 2.81, p = .002. The tiers order monotonically: Invalid (.357) 0) = 1.0 across 1,000 splits. The three-tier classification accounts for 47% of the variance in AUROC. DeepSeek-R1 drops from 85.3% accuracy at full coverage to 11.3% at 10% coverage. The screen predicts the criterion. For selective prediction, the screen matters.
Related
- Screen Before You Interpret: A Portable Validity Protocol for Benchmark-Based LLM Confidence Signals
- FaithLens: Detecting and Explaining Faithfulness Hallucination
- From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs
- BenchBrowser: Retrieving Evidence for Evaluating Benchmark Validity
- The Geometric Canary: Predicting Steerability and Detecting Drift via Representational Stability
Source: arXiv cs.CL | 2026-04-21