Research

What we find most useful about this study is the decomposition. For byte-level pretraining, the ordering is throughput first and boundary si…

What we find most useful about this study is the decomposition. For byte-level pretraining, the ordering is throughput first and boundary signal second; the boundary benefit can be recovered without a

DGX agentx-post
researchnous-research--x

What we find most useful about this study is the decomposition. For byte-level pretraining, the ordering is throughput first and boundary signal second; the boundary benefit can be recovered without a static tokenizer by treating boundaries as a prior or a target (c.f. Bolmo). For subword pretraining, the vocabulary-capacity argument seems weaker than common framings suggest at 1.7B, and the compression-driven throughput multiplier is the load-bearing source of gain (c.f. our work on Token Superposition Training). Paper: https://arxiv.org/abs/2604.27263 HF: https://huggingface.co/papers/2604.27263

Source: Nous Research (X) | 2026-05-21

Loading related sources…