Step-level Denoising-time Diffusion Alignment with Multiple Objectives
DGX agentarXiv:2604.14379v1 Announce Type: cross Abstract: Reinforcement learning (RL) has emerged as a powerful tool for aligning diffusion models with human preferences, typically by optimizing a single rewa