Research
What we find most useful about this study is the decomposition. For byte-level pretraining, the ordering is throughput first and boundary si…
What we find most useful about this study is the decomposition. For byte-level pretraining, the ordering is throughput first and boundary signal second; the boundary benefit can be recovered without a
What we find most useful about this study is the decomposition. For byte-level pretraining, the ordering is throughput first and boundary signal second; the boundary benefit can be recovered without a static tokenizer by treating boundaries as a prior or a target (c.f. Bolmo). For subword pretraining, the vocabulary-capacity argument seems weaker than common framings suggest at 1.7B, and the compression-driven throughput multiplier is the load-bearing source of gain (c.f. our work on Token Superposition Training). Paper: https://arxiv.org/abs/2604.27263 HF: https://huggingface.co/papers/2604.27263
Source: Nous Research (X) | 2026-05-21