Research
Towards joint scaling laws with optimal batch size schedules
arXiv:2607.27731v1 Announce Type: new Abstract: Modern deep learning typically keeps the batch size static throughout training, thus overlooking the joint effect of learning rate and batch size on the
arXiv:2607.27731v1 Announce Type: new Abstract: Modern deep learning typically keeps the batch size static throughout training, thus overlooking the joint effect of learning rate and batch size on the training dynamics. In this paper, we study the deep learning dynamics through the lens of convex optimization and derive a joint characterization of loss in terms of both schedules, applicable to general optimizers and model architectures. This characterization yields a closed-form optimal batch size schedule for any prescribed learning rate schedule, and further leads to joint scaling laws that consistently outperform static batch size baselines, highlighting the significance of dynamic batch size schedule in large language model training.
Related
- Accelerating SGDM via Learning Rate and Batch Size Schedules: A Lyapunov-Based Analysis
- Same Loss, Same Noise, Opposite Schedules: Noise Structure and Optimizer Normalization Jointly Determine Whether Learning-Rate Cooldown Helps
- Predictable Scaling Laws of Optimal Hyperparameters for LLM Continued Pre-training
- Practical Scaling Laws: Converting Compute into Performance in a Data-Constrained World
Source: arXiv cs.LG | 2026-07-31