Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning
DGX agentarXiv:2602.14868v2 Announce Type: replace-cross Abstract: Reinforcement learning has emerged as a powerful paradigm for unlocking reasoning capabilities in language models. However, relying on sparse