Model Releases

Dimensionality and Measurement Precision in HLE's Multiple-Choice Subset

arXiv:2607.27420v1 Announce Type: cross Abstract: Humanity's Last Exam (HLE) is widely used to evaluate frontier language models. HLE organizes its questions into eight subject-domain categories, whos

DGX agentpaper
model-releasesarxiv-cs-cl

arXiv:2607.27420v1 Announce Type: cross Abstract: Humanity's Last Exam (HLE) is widely used to evaluate frontier language models. HLE organizes its questions into eight subject-domain categories, whose subscores are often interpreted as evidence of distinct capabilities. However, no study has assessed whether these labels correspond to empirically separable latent constructs, nor whether the benchmark effectively differentiates between models of similar ability. We evaluate 29 LLMs on the text-only multiple-choice subset of HLE (J = 428 items) and apply psychometric methods to assess both the dimensionality of the benchmark and the distribution of its measurement precision. Fitting a two-parameter logistic IRT model, we find convergent evidence that HLE measures a single general reasoning factor: McDonald's omega_h = 0.998, domain labels explain only 3.5% of item response variance, within- and between-domain residual correlations are nearly identical (Cohen's d = 0.016), and domain-specific ability estimates are near-redundant with the total score (r geq 0.81). A separate analysis of the test information function reveals that measurement precision concentrates at moderate ability levels and drops sharply above heta = 0, where frontier models sit. These findings suggest that HLE's domain subscores do not warrant distinct capability interpretations and that the benchmark's ability to discriminate among the strongest models is limited.

Related

Source: arXiv cs.CL | 2026-07-31

Loading related sources…