Sample Where You Struggle: Sharpening Base Model Reasoning via Entropy-Guided Power Sampling
DGX agentarXiv:2606.09926v1 Announce Type: cross Abstract: Sampling from the sequence-level power distribution p^alpha elicits RL-level reasoning from base language models without any parameter updates, but th