Consolidation or Adaptation? PRISM: Disentangling SFT and RL Data via Gradient Concentration
DGX agentarXiv:2601.07224v2 Announce Type: replace Abstract: While Hybrid Supervised Fine-Tuning (SFT) followed by Reinforcement Learning (RL) has become the standard paradigm for training LLM agents, effectiv