A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models
arXiv:2607.26102v1 Announce Type: cross Abstract: Mathematical chain of thought (CoT) evaluation is commonly reduced to whether the final answer matches a reference. This conflates producing a correct