Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning

Published
Source
arXiv
Paper number
402
Field
Computer Vision
arXiv ID
2606.11683

Key points

  • It reformulates spatial reasoning as a two-stage process of hypothesis and verification rather than a single turn.
  • It strategically synthesizes complementary oblique views using VGGT monocular 3D reconstruction.
  • It is a training-free framework that can be applied at inference time without modifying the MLLM architecture.
  • It improves the VSI-Bench average by 5.2 points; at the sample level, the positive flip rate is 71.6% versus a negative flip rate of 28.4%, a 2.52 to 1 ratio.
  • On an A100 GPU, it takes about 11 seconds per sample; reducing the VGGT input to 20 frames preserves a 2.8-point gain in about 4 seconds.
  • Oblique Sweep outperforms bird's-eye views because it exposes spatial information while avoiding viewpoint-distribution mismatch.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)