EduArt: An educational-level benchmark for evaluating art history knowledge in large language models
arXiv:2607.02007v1 Announce Type: new Abstract: Large language models now score near ceiling on general benchmarks, but these aggregate measures reveal little about how models behave within single dis