REGEN: Replay-recycling for Expert-to-Generalist distillation with Offline Reinforcement Learning
DGX agentarXiv:2607.19450v1 Announce Type: cross Abstract: Large-scale online reinforcement learning (RL) is the predominant means of eliciting advanced abilities including long-term reasoning and agentic tool