Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key
DGX agentarXiv:2605.06638v2 Announce Type: replace Abstract: Reinforcement learning (RL) has been applied to improve large language model (LLM) reasoning, yet the systematic study of how training scales with t