Research

What we find most useful about TST is the decoupling. The training-time efficiency is fully separated from the inference-time architecture, …

What we find most useful about TST is the decoupling. The training-time efficiency is fully separated from the inference-time architecture, which makes TST a clean addition on top of other pretraining

DGX agentx-post
researchnous-research--x

What we find most useful about TST is the decoupling. The training-time efficiency is fully separated from the inference-time architecture, which makes TST a clean addition on top of other pretraining improvements: sparse attention, MoE routing, alternative tokenizers, optimizer changes. Most pretraining-efficiency interventions don't satisfy this criterion, and we think more should. TST is the first of several research releases from the Nous pretraining group going up over the next two weeks. If you want to work on problems like this, find us on Discord.

Source: Nous Research (X) | 2026-05-13

Loading related sources…