Research

Brilliant new paper from Meta. LLM judges get validated on accuracy against golden data. That says nothing about whether the verdict survive…

Brilliant new paper from Meta. LLM judges get validated on accuracy against golden data. That says nothing about whether the verdict survives when questioned. The Wiggle Framework stress-tests 9 front

DGX agentx-post
researchdair-ai--x

Brilliant new paper from Meta. LLM judges get validated on accuracy against golden data. That says nothing about whether the verdict survives when questioned. The Wiggle Framework stress-tests 9 frontier models across 14 judging tasks along three axes, stability under re-prompting, stability under a single challenge, and stability under sustained pressure. They find that every model wiggles. Verdicts flip 25 to 71% of the time under static pushback, and 62 to 91% against an adversarial persuader. Pressure that changes a judge's verdict is almost always net-corrupting against ground truth. Baseline jury majority strength turns out to be the best single-shot predictor of which items will move. Paper: https://arxiv.org/abs/2608.12645 Track more trending AI papers in our academy: https://academy.dair.ai/

Related

Source: DAIR.AI (X) | 2026-08-14

Loading related sources…