Research
PROOF-Gen: From Optimized Data to Better Distillation
Supervised fine-tuning on teacher-generated trajectories is the standard first stage for distilling tool-calling capabilities into deployable models. Post-training pipelines that drive shipped tool-ca
Supervised fine-tuning on teacher-generated trajectories is the standard first stage for distilling tool-calling capabilities into deployable models. Post-training pipelines that drive shipped tool-calling agents re-run this stage on a daily or weekly cadence, paying the frontier-teacher cost each cycle, yet the mechanism is generate-and-filter (keep the teacher’s passing trajectories, discard the rest) and each cycle leaves behind the same hard scenarios because failures supply no signal. On τ 2-bench, 57% of teacher trials fail, two-thirds of them near-misses (most tool calls correct, undone…
Related
- DomainPilot: Domain-Level Loss-Guided Two-Stage Data Mixture Optimization for Efficient Language Model Fine-Tuning
- Provenance-Grounded Gating and Adaptive Recovery in Synthetic Post-Training Data Curation
- Reasoning Quality Emerges Early: Data Curation for Reasoning Models
- Learning to Adapt SFT Data for Better Reasoning Generalization
Source: Apple ML Research | 2026-08-26