Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning
DGX agentarXiv:2512.11470v2 Announce Type: replace-cross Abstract: Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) dominate the post-training landscape for mathematical reasoning, yet differ funda