Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models
DGX agentarXiv:2508.04325v2 Announce Type: replace-cross Abstract: Large language models (LLMs) show significant potential in healthcare, prompting numerous benchmarks to evaluate their capabilities. However,