Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping
DGX agentarXiv:2605.10937v1 Announce Type: new Abstract: Recently, post-training methods based on reinforcement learning, with a particular focus on Group Relative Policy Optimization (GRPO), have emerged as t