How Reasoning Evolves from Post-Training Data: An Empirical Study Using Chess
DGX agentarXiv:2604.05134v2 Announce Type: replace Abstract: We study how reasoning evolves in a language model -- from supervised fine-tuning (SFT) to reinforcement learning (RL) -- by analyzing how a set of