How to Allocate Your Tokens? Scaling Laws with Training Steps and Batch Size
arXiv:2607.01487v1 Announce Type: new Abstract: We propose a scaling law that takes into account model size and training data while explicitly splitting the latter into training steps and batch size (