Safety
Do Not Step Into the Same River Twice: Learning to Reason from Trial and Error
arXiv:2510.26109v4 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) has significantly boosted the reasoning capability of language models (LMs). However, existing
arXiv:2510.26109v4 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) has significantly boosted the reasoning capability of language models (LMs). However, existing RLVR approaches train LMs based on their own on-policy responses and are constrained by the initial capability of LMs, thus prone to exploration stagnation, in which LMs fail to solve more training problems and cannot further learn from the training data. Some approaches try to address this by leveraging off-policy solutions to training problems, but rely on external expert guidance that is limited in availability and scalability. In this work, we propose LTE (Learning to reason from Trial and Error), an approach that hints LMs with their previously self-made mistakes, not requiring any external expert guidance. Experiments validate the effectiveness of LTE, which outperforms the normal group relative policy optimization (GRPO) by 5.02 in Pass@1 and 9.96 in Pass@k on average across six mathematical reasoning benchmarks for Qwen3-8B-Base and even performs better than methods that require external guidance. Further analysis confirms that LTE successfully mitigates exploration stagnation and enhances both exploitation and exploration during training. Our code is available at https://github.com/JamyDon/LTE.
Related
- RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
- Data-Efficient RLVR via Off-Policy Influence Guidance
- Think Outside the Policy: In-Context Steered Policy Optimization
- Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards
- From P(y|x) to P(y): Investigating Reinforcement Learning in Pre-train Space
- SSPO: Subsentence-level Policy Optimization
Source: arXiv cs.LG | 2026-04-17