Safety
CARE: Confidence-Aware Reasoning for Reliable Medical VQA
arXiv:2608.10964v1 Announce Type: cross Abstract: Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual qu
arXiv:2608.10964v1 Announce Type: cross Abstract: Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual question answering, yet these models suffer from extit{confidence miscalibration}---a systematic gap between expressed certainty and actual diagnostic accuracy that undermines clinical trust. We propose extbf{CARE}, a extbf{C}onfidence-extbf{A}ware medical extbf{RE}asoning framework that jointly optimizes accuracy and calibration through a dual-stage pipeline. First, a scalable Medical-CoT synthesis provides structured cold-start data for Supervised Fine-Tuning. Second, Group Relative Policy Optimization (GRPO) with a novel extbf{Confidence-Aware Reward (CAR)} mechanism ties the model's confidence to diagnostic correctness within the reward signal. Across three Medical VQA benchmarks, extbf{CARE} achieves the highest diagnostic accuracy while obtaining the lowest Expected Calibration Error and Hallucination Rate, establishing a foundation for trustworthy clinical decision support. Our code is available at https://github.com/anotherbricki/CARE.
Related
- Reinforcement-aware Knowledge Distillation for LLM Reasoning
- Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning
- Better Eyes, Better Thoughts: Why Vision Chain-of-Thought Fails in Medicine
- Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning
Source: arXiv cs.AI | 2026-08-12