Evaluating Research-Level Math Proofs via Strict Step-Level Verification
arXiv:2606.10799v1 Announce Type: new Abstract: Large Language Models (LLMs) struggle to rigorously verify complex mathematical proofs. Standard global evaluation approaches suffer from 'context poiso