VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning
- Published
- Source
- arXiv
- Paper number
- 1020
- Field
- Computer Vision
- arXiv ID
- 2608.26105
Key points
- It found that VLM-as-a-judge evaluations repeatedly fail through counting errors, disregard for fine-grained evidence, and misreading of rules, and replaced them with deterministic rule-based scores.
- A model trained on 300 procedurally generated tasks showed substantial transfer to 7 external benchmarks, including RISE-Video and MME-CoF-Pro (often +20 percentage points or more).
- Nearest-neighbor and counterfactual diagnostic experiments tested whether the gains came from genuine visual reasoning rather than memorization of instruction patterns.
- Comparing more than 30 image, video, and interleaved generators on the same task distribution showed that video is advantageous for spatiotemporal state tracking, while interleaving is advantageous for computational efficiency.
- RL with verifiable rewards raised the maze-navigation score from 0.2249 to 0.8692, and self-correction of intermediate answers was also observed.
Paper links
External research summaries. These are not HDATF publications or measured product results.