Applications
DataFoundry: Evolving Data Preparators via Recursive Self-Improvement
arXiv:2608.29966v1 Announce Type: new Abstract: Domain adaptation of large language models increasingly depends on constructing high-quality training data, yet existing data-preparation pipelines typi
arXiv:2608.29966v1 Announce Type: new Abstract: Domain adaptation of large language models increasingly depends on constructing high-quality training data, yet existing data-preparation pipelines typically address quality only after generation through post-hoc filtering. This creates a fundamental mismatch: data-quality issues often originate from the construction process itself, while quality control is applied only to its outputs. We introduce extsc{DataFoundry}, a framework for extbf{evolving data preparators through recursive self-improvement} before large-scale data production. extsc{DataFoundry} represents a data preparator as an evolvable runtime specification and instantiates its evolution with a extsc{Skills-as-Modules} architecture, in which a central extsc{Controller} orchestrates modular skills to compile executable runtimes, diagnose deficiencies on small pilot sets using domain-appropriate criteria, and translate diagnostic feedback into adapters that revise individual preparation components while preserving stable interfaces. We evaluate extsc{DataFoundry} on DataPrep-Bench across mathematics, finance, law, and medicine, and find that recursively evolved preparators produce training data with higher downstream utility than baselines. Experiments across different backbones further demonstrate that these improvements are not tied to a particular model, while analyses and case studies further reveal the framework's optimization dynamics and illustrate how its evolution unfolds in practice.
Related
- IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages
- Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training
- Learning Dynamic User Personas from Implicit Interaction Streams via Iterative Refinement
- FOCUS: Decoupling Expert Personas in LLMs to Enhance Domain Expert Capabilities
Source: arXiv cs.CL | 2026-09-01