Learning-Zone Energy: Online Data Selection for Efficient RL Post-Training
DGX agentarXiv:2605.17003v1 Announce Type: cross Abstract: Reinforcement Learning (RL) post-training has emerged as the dominant paradigm for eliciting mathematical reasoning in Large Language Models (LLMs), y