Model Releases
When Counterbalancing Hides the Bias: Access-Conditioned Position Lock in Forced-Choice LLM Evaluation
arXiv:2607.10202v2 Announce Type: replace Abstract: Forced-choice probes with counterbalanced orientations are a standard tool for measuring language-model 'value dispositions,' and a concentration/ex
arXiv:2607.10202v2 Announce Type: replace Abstract: Forced-choice probes with counterbalanced orientations are a standard tool for measuring language-model "value dispositions," and a concentration/extremity index over repeated draws is read as how sharply a model commits. We show this estimator is not identifiable at its low end: counterbalancing, meant to remove position bias, instead maps a position-lock (a model returning the same answer letter regardless of content) onto the same near-0.5 signature as genuine neutrality, so "softness" and "non-engagement" cannot be distinguished from a content-independent letter-bias. Across nine models the fraction of position-locked items tracks the index almost perfectly (r=-0.986) - a structural consequence, not a finding: the informative quantity is the residual from that bound, the commitment a model shows on the items it does engage. The three models the index reads as soft are the most position-locked, and the lock resolves under access configurations that permit reasoning (as-deployed CLI -> raw API -> reasoning-enabled), moving the index 0.06->0.63 and 0.34->0.66->0.64 while lock collapses (Opus: 0.61->0.39->0.22); this establishes the low readings as an identifiability failure, not a disposition. The reasoning-enabled reading is not a "true" value either; the point is that the index alone cannot identify commitment at its low end. Access configuration (deployment client, reasoning on/off) is one generator of this failure, shown on two access paths (an Anthropic subscription CLI and a DeepSeek client), and in a frozen benchmark it is confounded with provider. We contribute a position-lock diagnostic that must accompany any concentration reading, and show that a concentration-blind audit risks reporting neutrality where a reasoning-permitting condition yields concentrated choices. The direction-flip component, largely robust to the artifact, still identifies genuine cross-model disagreement.
Source: arXiv cs.LG | 2026-08-11