Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization
arXiv:2510.21583v2 Announce Type: replace Abstract: Recent Progress in post-training flow matching for text-to-image (T2I) generation with Group Relative Policy Optimization (GRPO) has demonstrated st