Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling
DGX agentarXiv:2602.11146v2 Announce Type: replace-cross Abstract: Preference optimization for diffusion and flow-matching models relies on reward functions that are both discriminatively robust and computatio