Safety
Product-of-Experts Training Reduces Dataset Artifacts in Natural Language Inference
arXiv:2604.19069v1 Announce Type: cross Abstract: Neural NLI models overfit dataset artifacts instead of truly reasoning. A hypothesis-only model gets 57.7% in SNLI, showing strong spurious correlatio
arXiv:2604.19069v1 Announce Type: cross Abstract: Neural NLI models overfit dataset artifacts instead of truly reasoning. A hypothesis-only model gets 57.7% in SNLI, showing strong spurious correlations, and 38.6% of the baseline errors are the result of these artifacts. We propose Product-of-Experts (PoE) training, which downweights examples where biased models are overconfident. PoE nearly preserves accuracy (89.10% vs. 89.30%) while cutting bias reliance by 4.71% (bias agreement 49.85% to 45%). An ablation finds lambda = 1.5 that best balances debiasing and accuracy. Behavioral tests still reveal issues with negation and numerical reasoning.
Related
- Fairness is Not Flat: Geometric Phase Transitions Against Shortcut Learning
- Reinforcement-aware Knowledge Distillation for LLM Reasoning
- Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
- Rethinking Token-Level Credit Assignment in RLVR: A Polarity-Entropy Analysis
Source: arXiv cs.AI | 2026-04-22