VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

Published
Source
arXiv
Paper number
1020
Field
Computer Vision
arXiv ID
2608.26105

Key points

  • It found that VLM-as-a-judge evaluations repeatedly fail through counting errors, disregard for fine-grained evidence, and misreading of rules, and replaced them with deterministic rule-based scores.
  • A model trained on 300 procedurally generated tasks showed substantial transfer to 7 external benchmarks, including RISE-Video and MME-CoF-Pro (often +20 percentage points or more).
  • Nearest-neighbor and counterfactual diagnostic experiments tested whether the gains came from genuine visual reasoning rather than memorization of instruction patterns.
  • Comparing more than 30 image, video, and interleaved generators on the same task distribution showed that video is advantageous for spatiotemporal state tracking, while interleaving is advantageous for computational efficiency.
  • RL with verifiable rewards raised the maze-navigation score from 0.2249 to 0.8692, and self-correction of intermediate answers was also observed.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)