DataPrep-Bench: Benchmarking LLMs as Training Data Preparators
DGX agentarXiv:2607.20465v1 Announce Type: cross Abstract: The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unified benchmark exists to measure how