Model Releases
Dimensionality and Measurement Precision in HLE's Multiple-Choice Subset
arXiv:2607.27420v1 Announce Type: cross Abstract: Humanity's Last Exam (HLE) is widely used to evaluate frontier language models. HLE organizes its questions into eight subject-domain categories, whos
arXiv:2607.27420v1 Announce Type: cross Abstract: Humanity's Last Exam (HLE) is widely used to evaluate frontier language models. HLE organizes its questions into eight subject-domain categories, whose subscores are often interpreted as evidence of distinct capabilities. However, no study has assessed whether these labels correspond to empirically separable latent constructs, nor whether the benchmark effectively differentiates between models of similar ability. We evaluate 29 LLMs on the text-only multiple-choice subset of HLE (J = 428 items) and apply psychometric methods to assess both the dimensionality of the benchmark and the distribution of its measurement precision. Fitting a two-parameter logistic IRT model, we find convergent evidence that HLE measures a single general reasoning factor: McDonald's omega_h = 0.998, domain labels explain only 3.5% of item response variance, within- and between-domain residual correlations are nearly identical (Cohen's d = 0.016), and domain-specific ability estimates are near-redundant with the total score (r geq 0.81). A separate analysis of the test information function reveals that measurement precision concentrates at moderate ability levels and drops sharply above heta = 0, where frontier models sit. These findings suggest that HLE's domain subscores do not warrant distinct capability interpretations and that the benchmark's ability to discriminate among the strongest models is limited.
Related
- The Blessing of Dimensionality: How Near-Orthogonality in High-Dimensional Spaces Explains Temporal Portability
- Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong
- Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets
- Domain Fine-Tuning vs. Retrieval-Augmented Generation for Medical Multiple-Choice Question Answering: A Controlled Comparison at the 4B-Parameter Scale
Source: arXiv cs.CL | 2026-07-31