MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference
DGX agentarXiv:2602.15206v2 Announce Type: replace Abstract: Reward learning typically relies on a single feedback type or combines multiple feedback types using manually weighted loss terms. Currently, it rem