Model Releases
DoublesEval: Diagnosing Multi-Agent Tactical Reasoning in Vision-Language Models via Professional Doubles Badminton
arXiv:2608.24439v1 Announce Type: new Abstract: Visual Language Models (VLMs) excel at describing visible scene content but struggle to reason about dynamic multi-agent interactions, where action sema
arXiv:2608.24439v1 Announce Type: new Abstract: Visual Language Models (VLMs) excel at describing visible scene content but struggle to reason about dynamic multi-agent interactions, where action semantics depend on coordinated roles and spatial-temporal dependencies. We formalize this capability as extbf{multi-agent tactical reasoning} and introduce extbf{DoublesEval}, a diagnostic evaluation framework that leverages professional doubles badminton as a structurally tractable testbed. DoublesEval employs a key-moment-based protocol that decomposes rallies into tactically salient instants and probes models across four interpretable dimensions: atomic recognition, intra-segment composite understanding, cross-segment causal reasoning, and high-level tactical abstraction. This design isolates where reasoning fails, rather than merely measuring answer correctness. To address observed failure modes, we propose extbf{TacticCheck}, a lightweight constraint-guided test-time consistency checker that reranks candidate answers using the model's own lower-level tactical predictions, requiring no parameter updates or ground-truth labels at inference time. Evaluating four representative open-source VLMs on 60 curated rallies (yielding sim9.6K structured instances) via a zero-shot protocol, we find that models remain weak across all diagnostic levels, with especially clear bottlenecks in spatial state, interaction binding, and terminal evidence. TacticCheck delivers consistent gains across all evaluated models, while still leaving a substantial gap to robust tactical reasoning. These results highlight the need for structured, interaction-aware evaluation paradigms for next-generation VLMs. The source code is available in href{https://github.com/Chengjt1999/DoublesEval}{extcolor{blue}{our GitHub repository}}.
Related
- ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?
- ChronoVision: Temporal Reasoning via Latent State Reconstruction
- 4DP-QA: Scalable QA for 4D Perception in Vision Language Models
Source: arXiv cs.CV | 2026-08-26