Learning from Saturated Data: Signals Beyond Correctness for LLM Training
arXiv:2606.01436v1 Announce Type: new Abstract: The growing capabilities of large language models (LLMs) have led to the saturation of many benchmarks and training datasets used to improve them. Motiv