Auditing Multimodal LLM Raters: Central Tendency Bias in Clinical Ordinal Scoring
arXiv:2605.16386v1 Announce Type: new Abstract: Multimodal large language models (LLMs) are increasingly explored as automated evaluators in clinical settings, yet their scoring behavior on ordinal cl