Hierarchical Denoising For Multi-Step Visual Reasoning
- Published
- Source
- arXiv
- Paper number
- 652
- Field
- Computer Vision
- arXiv ID
- 2607.15278
Key points
- We solve the limitations of both fast but dumb streaming and smart but slow bidirectional video methods with a tree-structured approach.
- Across six multi-step reasoning tasks, including maze solving, Sokoban, and Hanoi towers, success rises from 34.22 to 60.29, a 76.2% improvement.
- The streaming version runs at 0.70 seconds per frame, 54x faster than bidirectional diffusion, while still achieving higher reasoning accuracy.
- Using only 2% of the training data preserves 82.9% of the performance of the full data, showing strong data efficiency.
- Real robotic arm maze experiments also show robust success, confirming applicability in the physical world.
Paper links
External research summaries. These are not HDATF publications or measured product results.