Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks
arXiv:2608.02966v1 Announce Type: new Abstract: Most multiple-choice question (MCQ) benchmarks evaluate Large Language Models (LLMs) only by whether they select the correct answers. This binary scorin