Causal Direct Preference Optimization for Distributionally Robust Generative Recommendation
DGX agentarXiv:2603.22335v2 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO) guides large language models (LLMs) to generate recommendations aligned with user historical behavior dis