Research
Skaling: Chinchilla's Exponents Meet Kaplan's Coupling
arXiv:2608.07222v1 Announce Type: new Abstract: Neural scaling laws are foundational for language model development, yet standard formulations systematically under- and overestimate loss at data-scarc
arXiv:2608.07222v1 Announce Type: new Abstract: Neural scaling laws are foundational for language model development, yet standard formulations systematically under- and overestimate loss at data-scarce and overtraining extremes. This failure originates in the underlying assumption that model size and training data impact the loss independently. To address this, we introduce the Skaling law, a generalized functional form that couples model capacity and data through a single interaction exponent. This simple extension reduces the Mean Absolute Percentage Error (MAPE) by 1.5-3x across both interpolation and extrapolation regimes. When paired with a sparse grid strategy restricted to low-compute regimes, the Skaling law achieves accurate full-grid extrapolation using approximately 10x less compute than uniform sweeps. By enabling reliable performance prediction from small-scale experiments, the Skaling law provides a more robust and resource-efficient framework for allocating compute budgets in next-generation model training.
Related
- Neural Scaling Universality: If Exponents Are Fixed, Time to Understand Coefficients
- Emergent Capabilities Arise Randomly from Learning Sparse Attention Patterns
- InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and Repetition
- Is More Data Worth the Cost? Dataset Scaling Laws in a Tiny Attention-Only Decoder
Source: arXiv cs.CL | 2026-08-10