NebulaExp-8B: An Empirical Post-Training Pipeline via Full-Scale Ablation Research
DGX agentarXiv:2606.26671v1 Announce Type: new Abstract: Post-training alignment determines the reasoning and human preference following capabilities of large language models, yet most existing works withhold