Model Releases
Evaluation-as-Search: Adaptive Discovery of Grounding Failures in Meeting Assistants
arXiv:2608.20392v1 Announce Type: cross Abstract: LLM-powered meeting assistants are deployed at scale, yet systematic evaluation of their grounding fidelity remains limited to static benchmarks that
arXiv:2608.20392v1 Announce Type: cross Abstract: LLM-powered meeting assistants are deployed at scale, yet systematic evaluation of their grounding fidelity remains limited to static benchmarks that miss failure modes tied to specific discourse structures or reasoning demands. We propose Evaluation-as-Search (EaS), a feedback-driven methodology that frames quality evaluation as an adaptive search over the space of natural questions a meeting participant might ask. Rather than sampling uniformly, EaS learns from evaluator feedback across iterations to concentrate probing effort on cognitive demands where failures are most likely, guided by a UCB-scored coverage map and blind multi-dimensional quality evaluation. Using EaS, we construct MeetingProbe, a benchmark of over 3{,}000 annotated question--answer pairs spanning 20 transcripts from three meeting genres and three LLM assistants. In ablations, adaptive search surfaces 2.5imes more failures than random probing (7.1% vs. 2.9% finding rate), with the strategic planner contributing the largest individual effect. Across three models, we observe a clear capability gradient and identify eight recurring failure categories dominated by discourse-pragmatic challenges rather than factual recall errors. We further validate MeetingProbe across multiple model families and providers, finding a clean capability gradient and a curated subset of universal failures that no model handles. MeetingProbe is released publicly to support reproducible evaluation of meeting assistant grounding fidelity.
Related
- Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
- Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems
- Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic Critique
Source: arXiv cs.AI | 2026-08-24