PAFO: Pareto Fairness Optimization for Personalized Reward Modeling
DGX agentarXiv:2606.07988v1 Announce Type: new Abstract: Large language models (LLMs) increasingly rely on reward models to align their outputs with diverse user preferences. While personalized reward models a