Safety
Stable but Wrong: When Learning Stabilizes Away from the Truth
arXiv:2603.21491v2 Announce Type: replace Abstract: Stable training is often treated as evidence that learning is succeeding, but stability characterizes optimization behavior rather than correctness
arXiv:2603.21491v2 Announce Type: replace Abstract: Stable training is often treated as evidence that learning is succeeding, but stability characterizes optimization behavior rather than correctness relative to an external objective. We study what happens when the signal being optimized remains persistently biased. We define Stable but Wrong (SBW) as a learning state in which the learning process remains stable under a task-appropriate operational criterion while the learned outcome remains systematically displaced from an independently defined objective. A minimal strongly convex model shows that a persistent bias in the update direction can shift the unique convergence point away from the true optimum. Controlled experiments in reinforcement learning, supervised learning, and continual fine-tuning of a large language model reveal a recurring separation between apparently normal optimization and correctness under static and feedback-coupled biases. Recovery-stage clean-data access and exploration interventions further show that subsequent trajectories can remain modifiable, although the interventions differ in protocol and do not imply a shared mechanism. The results expose a basic limit of optimization stability as a reliability signal in persistent and feedback-coupled learning systems.
Source: arXiv cs.LG | 2026-08-28