Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling
DGX agentarXiv:2606.03102v1 Announce Type: new Abstract: Test-time scaling improves the reasoning performance of large language models but incurs substantial cost in both total computation and latency. Existin