Rethinking Local Learning: A Cheaper and Faster Recipe for LLM Post-Training
DGX agentarXiv:2605.04913v1 Announce Type: new Abstract: LLM post-training typically propagates task gradients through the full depth of the model. Although this end-to-end structure is simple and general, it