Tutorials
The second technique was two-phase post-training. We first trained purely for capability, then added a latency penalty calibrated from real …
The second technique was two-phase post-training. We first trained purely for capability, then added a latency penalty calibrated from real dogfooding data based on the CDF of how long users stay on S
The second technique was two-phase post-training. We first trained purely for capability, then added a latency penalty calibrated from real dogfooding data based on the CDF of how long users stay on SWE-check before switching off. Training capability and latency jointly from the start caused the model to collapse into shallow-but-fast local optima; separating the phases let it build real bug-detection skill first and then learn to compress it.
Related
- LLMs learn backwards, and the scaling hypothesis is bounded. [D]
- A Severity-Based Curriculum Learning Strategy for Arabic Medical Text Generation
- What I’ve been building: ATOM Report, post-training course, finishing my book, and ongoing research
- A Data-driven Loss Weighting Scheme across Heterogeneous Tasks for Image Denoising
Source: Cognition AI (X) | 2026-04-14