Beyond Mode Collapse: Distribution Matching for Diverse Reasoning
arXiv:2605.19461v1 Announce Type: new Abstract: On-policy reinforcement learning methods like GRPO suffer from mode collapse: they exhibit reduced solution diversity, concentrating probability mass on