DRM: Diffusion-based Reward Model With Step-wise Guidance
arXiv:2605.25661v2 Announce Type: replace Abstract: Current mainstream methods of aligning diffusion models with human preferences typically employ VLM-based reward models. However, these reward model