Model Releases
ProactBench: Beyond What The User Asked For
arXiv:2605.09228v1 Announce Type: cross Abstract: Most LLM benchmarks score how well a model responds to explicit requests. They leave unmeasured a different conversational ability: noticing and actin
arXiv:2605.09228v1 Announce Type: cross Abstract: Most LLM benchmarks score how well a model responds to explicit requests. They leave unmeasured a different conversational ability: noticing and acting on needs the user has implied but not said. We call this conversational proactivity. ProactBench decomposes it into three phase-tied types: extsc{Emergent}, inference from a single disclosed anchor; extsc{Critical}, synthesis across multiple anchors; and extsc{Recovery}, grounded forward-looking value after task completion. We operationalise the benchmark with three agents: a Planner, a User Agent, and an Assistant Model. Their information asymmetries defend against style-confounded scoring, rubric leakage, external-context contamination, and information dumps. The released corpus contains 198 curated dialogues with 624 trigger points across 24 communication styles drawn from a psychometric inventory and audited by an independent LLM judge. Across 16 frontier and open-weight models, extsc{Recovery} is both difficult and weakly predicted by six standard benchmarks, making it a useful new evaluation signal.
Source: arXiv cs.AI | 2026-05-12