Research
Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning
arXiv:2607.19790v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has substantially improved language-model reasoning, yet its extension to vision-language models r
arXiv:2607.19790v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has substantially improved language-model reasoning, yet its extension to vision-language models remains constrained by the lack of training data that are simultaneously broad, exactly verifiable, and reproducible. We introduce Trace, a taxonomy-guided environment for multidomain visual reasoning. Trace factorizes task construction into a scene grammar and an executable task program, separating visual realization from answer computation. A shared semantic state determines the rendered image, prompt, typed answer, verifier state, and replayable instance trace. The resulting environment comprises 1,000 tasks over 277 scene grammars and 11 visual domains, with controlled semantic and visual variation. RLVR on 64,000 Trace instances improves the macro-average across 24 external benchmarks by 3.51 percentage points for Qwen2.5-VL-3B and 4.06 points for Qwen2.5-VL-7B, providing evidence that broad procedural training can transfer beyond the generated task distributions. Project page: https://maveryn.github.io/trace/.
Related
- When Does a Video-Language Model Stop Watching? Reward Strength Controls the Formation and Reversal of Visual Shortcuts in Multimodal RLVR
- Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention
- Tandem Reinforcement Learning with Verifiable Rewards
- CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning
- Reasoning Arena: Trace Tournaments When Verifiable Rewards Fall Short
Source: arXiv cs.CV | 2026-07-23