Surflo: Consistent 3D Surface Flow Model with Global State

Published
Source
arXiv
Paper number
414
Field
Computer Vision
arXiv ID
2606.13644

Key points

  • It compresses variable-view input into K latent tokens using a frozen VGGT backbone and Perceiver compression, producing a fixed-size representation independent of the number of views.
  • A flow-matching decoder transports each point from noise to surface independently, which supports arbitrary resolution from thousands to one million points.
  • Inference-time photometric and monodepth guidance resolves local inconsistencies that arise from independent decoding.
  • It is the only feed-forward method that overcomes the limitations of both per-view pointmap duplication and fixed-resolution global-latent approaches.
  • On eight 3D reconstruction benchmarks, it reaches the top feed-forward level and is more than 10x faster than optimization-based methods.
  • It plans to build and release a dataset that adds watertight meshes to about 10.5K DL3DV scenes.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)