Demystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining
arXiv:2606.16246v2 Announce Type: replace Abstract: As AI labs approach a data ceiling where compute capacity outpaces the rate of new high-quality text generation, language model pretraining is shift