Dynamic-TreeRPO: Breaking the Independent Trajectory Bottleneck with Structured Sampling
DGX agentarXiv:2509.23352v3 Announce Type: replace-cross Abstract: The integration of Reinforcement Learning (RL) into flow matching models for text-to-image (T2I) generation has driven substantial advances in