Glob3R: Global Structure-from-Motion with 3D Foundation Models

Published
Source
arXiv
Paper number
599
Field
Computer Vision
arXiv ID
2607.09225

Key points

  • It uses a frozen 3D foundation model, Pi3X, plus a lightweight dense matching head to turn feed-forward prediction into an optimizable geometric constraint.
  • A keyframe-based sliding-window strategy propagates tracks and relative poses across windows, which overcomes chunk-level limits.
  • On KITTI, it achieves a mean RMSE of 13.21 m, which is a 10 to 50 percent improvement over recent methods such as Scal3R at 14.55 and LoGeR at 18.65.
  • On ETH3D, it achieves R@1 of 91.58 and T@1 of 73.27, which substantially improves rotation and translation accuracy over the latest learning-based SfM method, AMB3R.
  • It improves PSNR by 2 to 3 dB in NeRF and neural rendering, which demonstrates the practical value of refined poses.
  • Global motion averaging and bundle adjustment resolve scale mismatch and recover dense geometry.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)