Plan-and-Verify Video Reward Reasoning with Spatio-Temporal Scene Graph Grounding
DGX agentarXiv:2606.11838v1 Announce Type: new Abstract: Reward models for text-to-video (T2V) generation guide post-training but often fail at fine-grained semantic alignment. We trace this to two structural