DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data
arXiv:2607.24717v1 Announce Type: cross Abstract: Pretraining data processing is critical to the downstream performance of Large Language Models (LLMs). However, many existing approaches define a fixe