Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence Attribution
DGX agentarXiv:2605.19228v1 Announce Type: cross Abstract: Large Language Models have achieved strong performance on reasoning tasks with objective answers by generating step-by-step solutions, but diagnosing