StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement Learning
DGX agentarXiv:2605.11922v1 Announce Type: cross Abstract: Existing code reasoning methods primarily supervise final code outputs, ignoring intermediate states, often leading to reward hacking where correct an