SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation
DGX agentarXiv:2608.09271v1 Announce Type: cross Abstract: Group-based reinforcement learning objectives such as GRPO can allocate learning signal poorly across prompt difficulty: under binary rewards, group n