Agents

There is much less signal in agent leaderboards than the rankings imply. A four-facet Generalizability Theory decomposition across TheAgentC…

There is much less signal in agent leaderboards than the rankings imply. A four-facet Generalizability Theory decomposition across TheAgentCompany, tau-squared-bench, and AppWorld finds the agent main

DGX agentx-post
agentsdair-ai--x

There is much less signal in agent leaderboards than the rankings imply. A four-facet Generalizability Theory decomposition across TheAgentCompany, tau-squared-bench, and AppWorld finds the agent main effect accounts for under 3% of total variance in every dataset and check type. The agent-by-task interaction accounts for 7 to 23%. Leaderboards are ranking specialization. Three estimators agree to three decimal places, Henderson Method-I, REML via lme4, and a Bayesian binomial GLMM. Aggregate reliability collapses where deployment actually hurts. On the hardest task quartile, reliability on tau-squared action checks falls from 0.752 to 0.000. Training-cell reliability also correlates negatively with held-out reliability at minus 0.90, so the designs that look most reliable replicate worst. Population-level diagnostics do transfer, with the capability-gap ratio stable at 0.35 to 0.40 across enterprise benchmarks. Per-family rankings invert. Paper: https://arxiv.org/abs/2608.11323 Track more trending AI papers in our academy: https://academy.dair.ai/

Related

Source: DAIR.AI (X) | 2026-08-13

Loading related sources…