Agents
There is much less signal in agent leaderboards than the rankings imply. A four-facet Generalizability Theory decomposition across TheAgentC…
There is much less signal in agent leaderboards than the rankings imply. A four-facet Generalizability Theory decomposition across TheAgentCompany, tau-squared-bench, and AppWorld finds the agent main
There is much less signal in agent leaderboards than the rankings imply. A four-facet Generalizability Theory decomposition across TheAgentCompany, tau-squared-bench, and AppWorld finds the agent main effect accounts for under 3% of total variance in every dataset and check type. The agent-by-task interaction accounts for 7 to 23%. Leaderboards are ranking specialization. Three estimators agree to three decimal places, Henderson Method-I, REML via lme4, and a Bayesian binomial GLMM. Aggregate reliability collapses where deployment actually hurts. On the hardest task quartile, reliability on tau-squared action checks falls from 0.752 to 0.000. Training-cell reliability also correlates negatively with held-out reliability at minus 0.90, so the designs that look most reliable replicate worst. Population-level diagnostics do transfer, with the capability-gap ratio stable at 0.35 to 0.40 across enterprise benchmarks. Per-family rankings invert. Paper: https://arxiv.org/abs/2608.11323 Track more trending AI papers in our academy: https://academy.dair.ai/
Related
- // Adapt the Interface, Not the Model // I am fascinated by the results across my cheap-model-plus-good-harness builds. This new paper also …
- i haven't seen a model that just works across agent harnesses. seems like it should exist. great opportunity for open-weight models. any tho…
- Highly recommended. I've often claimed there's huge alpha in building agent harnesses. Turns out harnesses are compositional generalizers. T…
Source: DAIR.AI (X) | 2026-08-13