Unlocking the Working Memory of Large Language Models for Latent Reasoning
DGX agentarXiv:2605.30343v1 Announce Type: cross Abstract: To improve the reasoning capabilities of large language models, test-time compute is typically scaled by generating intermediate tokens before the fin