Cost-Efficient Estimation of General Abilities Across Benchmarks
DGX agentarXiv:2604.01418v2 Announce Type: replace Abstract: Thousands of diverse benchmarks have been developed to measure the quality of large language models (LLMs). Yet prior work has demonstrated that LLM