Mitigating Distribution Sharpening in Math RLVR via Distribution-Aligned Hint Synthesis and Backward Hint Annealing
DGX agentarXiv:2604.07747v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) can improve low-k reasoning accuracy while narrowing solution coverage on challenging math que