Generating Pretraining Tokens from Organic Data for Data-Bound Scaling
arXiv:2605.17849v1 Announce Type: cross Abstract: LLM pretraining is shifting from a compute-bound to a data-bound regime, where available human (organic) text falls far short of scaling demands. Howe