Open-World Evaluations for Measuring Frontier AI Capabilities
DGX agentarXiv:2605.20520v1 Announce Type: new Abstract: Benchmark-based evaluation remains important for tracking frontier AI progress. But it can both overstate and understate deployed capability because it