UniVR: Thinking in Visual Space for Unified Visual Reasoning

Published
Source
arXiv
Paper number
629
Field
Computer Vision
arXiv ID
2607.12800

Key points

  • VR-GRPO combines overall outcome rewards with rewards for error-prone steps to learn logical and physical consistency in visual reasoning.
  • It built the VR-X benchmark from 16 sources, bringing long-horizon manipulation, spatial puzzles, and physical reasoning together in a purely visual format.
  • UniVR improved by up to 25% on VR-X, and stronger visual reasoning also improved performance on multiple multimodal-understanding benchmarks.
  • It released code, data, and models for research on planning and reasoning in visual space without textual explanations.
  • The 34B model and long-video training require substantial computational resources, and step rewards also depend on a general-purpose VLM that lacks detailed physical knowledge.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)