Model Releases
Risk reversal for least squares estimators under nested convex constraints
arXiv:2601.16041v2 Announce Type: replace-cross Abstract: In constrained stochastic optimization, one expects that restricting the feasible set, provided it still contains the true parameter, should n
arXiv:2601.16041v2 Announce Type: replace-cross Abstract: In constrained stochastic optimization, one expects that restricting the feasible set, provided it still contains the true parameter, should not increase the statistical risk of the corresponding projection estimator. We show that this intuition can fail, even in basic settings. We investigate this phenomenon in the Gaussian sequence model. Given a compact, convex set Theta subseteq R^d, one observes [ Y = heta^star + sigma Z, qquad Z sim N(0, I_d), ] and seeks to estimate an unknown heta^star in Theta. Here, the maximum likelihood estimator over Theta coincides with the least squares estimator (LSE), given by the Euclidean projection of Y onto Theta. We construct an explicit example exhibiting risk reversal: for sufficiently large noise, there exist nested compact convex sets Theta_S subsetneq Theta_L and heta^star in Theta_S such that the LSE constrained to Theta_S has strictly larger squared-error risk than the LSE constrained to Theta_L. Moreover, we demonstrate that risk reversal can persist at the level of worst-case risk. Finally, we show that the phenomenon is not specific to Gaussian noise or squared-error risk, extending our results beyond both settings. We clarify this phenomenon by contrasting noise regimes. In the vanishing-noise limit, the risk is governed at first order by the statistical dimension of the tangent cone, and risk reversal cannot occur at the leading sigma^2 scale. In the diverging-noise regime, the risk instead depends on the global geometry of the constraint sets, and the embedding of Theta_S within Theta_L can reverse the risk ordering. These results reveal a previously unrecognized failure mode of the LSE. They demonstrate that in sufficiently noisy settings, tightening a constraint can paradoxically degrade statistical performance.
Source: arXiv cs.LG | 2026-07-28