Safety
DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling
arXiv:2604.19544v1 Announce Type: new Abstract: Multimodal reward models (MRMs) play a crucial role in aligning Multimodal Large Language Models (MLLMs) with human preferences. Training a good MRM req
arXiv:2604.19544v1 Announce Type: new Abstract: Multimodal reward models (MRMs) play a crucial role in aligning Multimodal Large Language Models (MLLMs) with human preferences. Training a good MRM requires high-quality multimodal preference data. However, existing preference datasets face three key challenges: lack of granularity in preference strength, textual style bias, and unreliable preference signals. Besides, existing open-source multimodal preference datasets suffer from substantial noise, yet there is a lack of effective and scalable curation methods to enhance their quality. To address these limitations, we propose extbf{DT2IT-MRM}, which integrates a extbf{D}ebiased preference construction pipeline, a novel reformulation of text-to-image (extbf{T2I}) preference data, and an extbf{I}terative extbf{T}raining framework that curates existing multimodal preference datasets for extbf{M}ultimodal extbf{R}eward extbf{M}odeling. Our experimental results show that DT2IT-MRM achieves new extbf{state-of-the-art} overall performance on three major benchmarks: VL-RewardBench, Multimodal RewardBench, and MM-RLHF-RewardBench.
Related
- ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System
- ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety
- Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs
- Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
- ProRe: A Proactive Reward System for GUI Agents via Reasoner-Actor Collaboration
Source: arXiv cs.AI | 2026-04-22