Model Releases
New research from Meta. (bookmark it) Most factuality work checks whether the claims in an answer are correct. GAMUT goes after the harder q…
New research from Meta. (bookmark it) Most factuality work checks whether the claims in an answer are correct. GAMUT goes after the harder question of whether the answer covers everything it should. I
New research from Meta. (bookmark it) Most factuality work checks whether the claims in an answer are correct. GAMUT goes after the harder question of whether the answer covers everything it should. It involves a two-level meta-rubric: It encodes content organization and importance, open-ended sets, ordered processes, and relationships between facts, then compiles mechanically into a flat checklist of binary rubrics an LLM judge can grade reliably. The benchmark itself is 1,813 questions grounded in real wearable imagery across 10 domains, each with an expert-verified rubric, plus a text-only variant. Across 14 frontier and open models the best score is only 58.7%, and results hold up regardless of which judge model grades. Coverage is where answers actually fail users, and the rubric-compilation recipe transfers to any task with hierarchical requirements. Paper: https://arxiv.org/abs/2607.19322 Learn to build effective AI agents in our academy: https://academy.dair.ai/
Related
- Qwen publishes new work on RL coding agents. (bookmark it) The idea is to continually build a verification system that co-evolves with AI ag…
- Agent evals are drifting away from production reality. Most benchmarks use clean tasks, well-specified requirements, deterministic metrics, …
- How far are we from agents that can self-generate world knowledge? The work proposes an outcome-based reward that measures how much an agent…
Source: DAIR.AI (X) | 2026-07-22