Forecasting Downstream Performance of LLMs With Proxy Metrics
DGX agentarXiv:2605.18607v1 Announce Type: new Abstract: Progress in language model development is often driven by comparative decisions: which architecture to adopt, which pretraining corpus to use, or which