Safety
Distribution-Free Conformal Prediction for Steel Fatigue Strength: Marginal Validity Is Not Enough
arXiv:2608.07589v1 Announce Type: cross Abstract: Predicting fatigue failure in steel components experimentally is costly because it requires testing across multiple compositions and processing condit
arXiv:2608.07589v1 Announce Type: cross Abstract: Predicting fatigue failure in steel components experimentally is costly because it requires testing across multiple compositions and processing conditions. This has spurred research on data-driven prediction models. Studies using the NIMS MatNavi steel fatigue dataset often report high point-prediction accuracy but rely on aggregate error metrics, leaving uncertainty about the reliability of individual predictions and whether accuracy is consistent across the fatigue-strength spectrum. This paper is the first to apply conformal prediction to steel fatigue strength, comparing five interval-construction methods across 50 independent data splits and distinguishing marginal coverage from coverage within specific sub-regions of the predicted property. A gradient-boosting point model achieves an R^2 of 0.976 +/- 0.009 and a mean absolute error of 18.3 +/- 2.3 MPa. Split-conformal prediction provides valid marginal coverage (0.918) but drops to 0.755 in the highest-strength quartile, where design margins are most critical, a pattern also observed with a Gaussian process baseline. A cross-fitted, normalized conformal method restores near-uniform coverage across all quartiles (0.869-0.938) without a significant increase in interval width, by scaling the interval based on a cross-fitted estimate of local prediction difficulty rather than using a single global width. Diagnostic analysis traces the residual gap in the highest-strength quartile to elevated residual variance (2.7x the pooled Q1-Q3 level) rather than a systematic bias, situating the shortfall against a proven distribution-free limit on exact conditional coverage. Marginal coverage claims for ML-based fatigue-strength predictions can conceal systematic unreliability precisely where engineering decisions are most risky; therefore, conditional coverage should be routinely assessed alongside marginal coverage.
Source: arXiv cs.LG | 2026-08-11