Surflo: Consistent 3D Surface Flow Model with Global State
- Published
- Source
- arXiv
- Paper number
- 414
- Field
- Computer Vision
- arXiv ID
- 2606.13644
Key points
- It compresses variable-view input into K latent tokens using a frozen VGGT backbone and Perceiver compression, producing a fixed-size representation independent of the number of views.
- A flow-matching decoder transports each point from noise to surface independently, which supports arbitrary resolution from thousands to one million points.
- Inference-time photometric and monodepth guidance resolves local inconsistencies that arise from independent decoding.
- It is the only feed-forward method that overcomes the limitations of both per-view pointmap duplication and fixed-resolution global-latent approaches.
- On eight 3D reconstruction benchmarks, it reaches the top feed-forward level and is more than 10x faster than optimization-based methods.
- It plans to build and release a dataset that adds watertight meshes to about 10.5K DL3DV scenes.
Paper links
External research summaries. These are not HDATF publications or measured product results.