SM4RT: Learning Structured Motion Geometry for 4D Reconstruction

Published
Source
arXiv
Paper number
720
Field
Computer Vision
arXiv ID
2607.22534

Key points

  • It proposes Structure-of-Motion, or SoM, a representation that expresses video motion as a combination of rigid motion bases rather than individual points.
  • It infers 3D geometry, motion, and scene motion structure in a single forward pass.
  • On the Kubric benchmark, it achieves state-of-the-art tracking accuracy, APD 90.62, and shows strong structure preservation on the deformation metric VM.
  • The DA3 backbone outperforms VGGT by a wide margin, and simple linear fusion is the most effective feature fusion method.
  • It is efficient at inference, taking 2.86 seconds per sample and using only 15.72 GB of GPU memory.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)