Safety
Penalizing Length: Uncovering Systematic Bias in Quality Estimation Metrics
arXiv:2510.22028v4 Announce Type: replace Abstract: Quality Estimation (QE) metrics are vital in machine translation for reference-free evaluation and increasingly serve as selection criteria in data
arXiv:2510.22028v4 Announce Type: replace Abstract: Quality Estimation (QE) metrics are vital in machine translation for reference-free evaluation and increasingly serve as selection criteria in data filtering and candidate reranking. However, the prevalence and impact of length bias in QE metrics have been underexplored. Through a systematic study of top-performing learned and LLM-as-a-Judge QE metrics across 10 diverse language pairs, we reveal two critical length effects: First, QE metrics consistently over-predict errors with increasing translation length, even for high-quality, error-free texts. Second, when candidates of comparable quality are available for the same source text, learned QE metrics tend to favor shorter translations, whereas LLM-as-a-Judge metrics range from being approximately length-neutral to favoring longer translations. These metric-dependent length effects can favor or penalize translations based on length rather than quality and can propagate into downstream pipelines that rely on QE signals for data selection or system optimization. We trace the root cause of MetricX-24 QE's cumulative-length penalty to skewed supervision distributions, in which longer error-free examples are underrepresented in training data. As a diagnostic intervention, we apply length normalization during training and show that this simple modification effectively decouples error prediction from sequence length, yielding more reliable QE signals across translations of varying length.
Source: arXiv cs.CL | 2026-08-11