Reassessing High-Performing LLMs on Polish Medical Exams: True Competence or Bias-Driven Performance?
DGX agentarXiv:2606.12250v1 Announce Type: new Abstract: Large language models (LLMs) in medicine are mainly evaluated using multiple-choice question answering (MCQA), which can overestimate real clinical abil