Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data
DGX agentarXiv:2410.13341v4 Announce Type: replace Abstract: High quality annotations are increasingly a bottleneck in the explosively growing machine learning ecosystem. Scalable evaluation methods that avoid