Model Releases
Mitigating Distribution Sharpening in Math RLVR via Distribution-Aligned Hint Synthesis and Backward Hint Annealing
arXiv:2604.07747v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) can improve low-k reasoning accuracy while narrowing solution coverage on challenging math que
arXiv:2604.07747v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) can improve low-k reasoning accuracy while narrowing solution coverage on challenging math questions, and pass@1 gains do not necessarily translate into better large-k performance. Existing hint-based approaches can make challenging questions trainable, but they leave two issues underexplored: teacher-student distribution mismatch and the need to reduce hint exposure to match no-hint evaluation. We address these issues through two components. Distribution-Aligned Hint Synthesis (DAHS) constructs verified teacher hints conditioned on student-style responses. Backward Hint Annealing (BHA) anneals hint exposure across difficulty buckets and uses per-question hint dropout to preserve no-hint updates throughout RL training. We evaluate the method in math RLVR under the DAPO training framework across AIME24, AIME25, and AIME26 using exttt{Qwen3-1.7B-Base} and exttt{Llama-3.2-1B-Instruct}. On exttt{Qwen3-1.7B-Base}, our method improves both pass@1 and pass@2048 relative to DAPO across the three AIME benchmarks. On exttt{Llama-3.2-1B-Instruct}, the gains are concentrated in the large-k regime. These results suggest that, in math RLVR, hint scaffolding is effective when it restores learnable updates on challenging questions early in training and is then gradually removed before no-hint evaluation.
Related
- Not All Tokens See Equally: Perception-Grounded Policy Optimization for Large Vision-Language Models
- When to Trust Tools? Adaptive Tool Trust Calibration For Tool-Integrated Math Reasoning
- TEMPER: Testing Emotional Perturbation in Quantitative Reasoning
- Don't Overthink It: Inter-Rollout Action Agreement as a Free Adaptive-Compute Signal for LLM Agents
- Beyond Social Pressure: Benchmarking Epistemic Attack in Large Language Models
Source: arXiv cs.CL | 2026-04-10