When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion
DGX agentarXiv:2607.20543v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) can improve one-sample accuracy while making a model worse under repeated sampling. We study thi