Model Releases
// Your LLM judge disagrees with the experts // LLM Judges can be tricky to build. Here is an interesting showcasing why: There propose a re…
// Your LLM judge disagrees with the experts // LLM Judges can be tricky to build. Here is an interesting showcasing why: There propose a reference-full benchmark of hundreds of complete human-to-huma
// Your LLM judge disagrees with the experts // LLM Judges can be tricky to build. Here is an interesting showcasing why: There propose a reference-full benchmark of hundreds of complete human-to-human dialogues written by professional script writers, with realistic turn densities and more than 36,000 per-turn human annotations across over 30,000 expert-generated turns. Conversational evaluation frameworks were mostly built for summarization, translation and short-form QA, and the metrics themselves are often derived and validated on synthetic data rather than human dialogue. Tested against expert judgment at this scale, both classical automatic metrics and reference-free LLM-as-a-judge approaches turn out to be unreliable. Their Mixture-of-Judges framework combines multiple evaluative signals and recovers roughly 30 percent better correlation with human assessment. Paper: https://arxiv.org/abs/2608.26131 Chat with Paper: https://academy.dair.ai/papers/evaluating-language-models-in-realistic-conversational-contexts-2608.26131
Related
- Here is my other heavily used pattern. Evaluator/Judge: Fable 5 Executor: GPT-5.5 I no longer wait for frontier models or am loyal to any. I…
- Interesting new research from IBM. If you pick models from benchmark deltas, some of that delta belongs to the phrasing rather than the mode…
- // Agents' Last Exam // Agents' Last Exam is a living benchmark of over 1,000 economically valuable tasks, built with 250+ industry experts …
- Another fascinating paper on LLM Judges. (bookmark it) It's from Amazon, and they show that if you run panels of LLM judges, averaging their…
Source: DAIR.AI (X) | 2026-08-30