Research
Weighting What Matters: Boosting Sample Efficiency in Medical Report Generation via Token Reweighting
arXiv:2604.21082v1 Announce Type: new Abstract: Training vision-language models (VLMs) for medical report generation is often hindered by the scarcity of high-quality annotated data. This work evaluat
arXiv:2604.21082v1 Announce Type: new Abstract: Training vision-language models (VLMs) for medical report generation is often hindered by the scarcity of high-quality annotated data. This work evaluates the use of a weighted loss function to improve data efficiency. Compared to standard cross-entropy loss, which treats all token prediction errors equally, the reweighted loss shifts the focus to semantically salient tokens with outsized clinical importance. In experiments on ophthalmological report generation, we show that this simple method improves efficiency across multiple data scales, achieving similar report quality with up to ten times less training data.
Related
- UIPress: Bringing Optical Token Compression to UI-to-Code Generation
- Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations
- VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors
- Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques, and Prospects
- When More Words Say Less: Decoupling Length and Specificity in Image Description Evaluation
Source: arXiv cs.CL | 2026-04-24