Research
PushupBench: Your VLM is not good at counting pushups
arXiv:2604.23407v1 Announce Type: cross Abstract: Large vision-language models (VLMs) can recognize extit{what} happens in video but fail to count extit{how many} times. We introduce extbf{PushupBench
arXiv:2604.23407v1 Announce Type: cross Abstract: Large vision-language models (VLMs) can recognize extit{what} happens in video but fail to count extit{how many} times. We introduce extbf{PushupBench}, 446 long-form clips (avg. 36.7s) for evaluating repetition counting. The best frontier model achieves 42.1% exact accuracy; open-source 4B models score sim6%, matching supervised baselines. We show that accuracy alone misleads -- weaker models exploit the modal count rather than reason temporally. Fine-tuning on counting with 1k samples transfers to general video understanding: MVBench (+2.15), PerceptionTest (+1.88), TVBench (+4.54), suggesting counting is a proxy for broader temporal reasoning.PushupBench incorporated in exttt{lmms-eval} (https://github.com/EvolvingLMMs-Lab/lmms-eval/pull/1262) and hosted on (pushupbench.com/)
Related
- CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception
- VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
- Distorted or Fabricated? A Survey on Hallucination in Video LLMs
Source: arXiv cs.AI | 2026-04-28