Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning
DGX agentarXiv:2602.01058v2 Announce Type: replace-cross Abstract: Post-training of reasoning LLMs is a holistic process that typically consists of an offline SFT stage followed by an online reinforcement lear