Model Releases

New research from Meta. (bookmark it) Most factuality work checks whether the claims in an answer are correct. GAMUT goes after the harder q…

New research from Meta. (bookmark it) Most factuality work checks whether the claims in an answer are correct. GAMUT goes after the harder question of whether the answer covers everything it should. I

DGX agentx-post
model-releasesdair-ai--x

New research from Meta. (bookmark it) Most factuality work checks whether the claims in an answer are correct. GAMUT goes after the harder question of whether the answer covers everything it should. It involves a two-level meta-rubric: It encodes content organization and importance, open-ended sets, ordered processes, and relationships between facts, then compiles mechanically into a flat checklist of binary rubrics an LLM judge can grade reliably. The benchmark itself is 1,813 questions grounded in real wearable imagery across 10 domains, each with an expert-verified rubric, plus a text-only variant. Across 14 frontier and open models the best score is only 58.7%, and results hold up regardless of which judge model grades. Coverage is where answers actually fail users, and the rubric-compilation recipe transfers to any task with hierarchical requirements. Paper: https://arxiv.org/abs/2607.19322 Learn to build effective AI agents in our academy: https://academy.dair.ai/

Related

Source: DAIR.AI (X) | 2026-07-22

Loading related sources…