Woodpecker Distillation: Weak Models Diagnose Reasoning Bugs in Strong Models
arXiv:2608.05168v1 Announce Type: new Abstract: Large language models often fail on reasoning tasks despite possessing the capability to solve them. We argue that many such failures arise from localiz