SLPO: Scaling Latent Reasoning via a Surrogate Policy
DGX agentarXiv:2607.19691v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoner