Hierarchical Denoising For Multi-Step Visual Reasoning

Published
Source
arXiv
Paper number
652
Field
Computer Vision
arXiv ID
2607.15278

Key points

  • We solve the limitations of both fast but dumb streaming and smart but slow bidirectional video methods with a tree-structured approach.
  • Across six multi-step reasoning tasks, including maze solving, Sokoban, and Hanoi towers, success rises from 34.22 to 60.29, a 76.2% improvement.
  • The streaming version runs at 0.70 seconds per frame, 54x faster than bidirectional diffusion, while still achieving higher reasoning accuracy.
  • Using only 2% of the training data preserves 82.9% of the performance of the full data, showing strong data efficiency.
  • Real robotic arm maze experiments also show robust success, confirming applicability in the physical world.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)