World from Motion: Generative Dynamic Gaussian Reconstruction from Monocular Video

Published
Source
arXiv
Paper number
547
Field
Computer Vision
arXiv ID
2607.01202

Key points

  • It is the first method to condition a video generation model on a dynamic 3DGS representation, rendering geometry, appearance, and 3D motion according to reference and target views.
  • It achieves state-of-the-art results on the DyCheck benchmark with mPSNR 19.98, mSSIM 0.716, and mLPIPS 0.178.
  • In the optimization step that distills generated samples back into dynamic 3DGS, it also improves 3D tracking accuracy, with PCK@0.05 rising from 0.824 for MoSca to 0.862.
  • Reconstruction quality improves monotonically as the number of virtual cameras increases and converges after 8 cameras, with mPSNR rising from 18.69 to 19.78.
  • It outperforms both generation-based methods such as CAT4D and Vista4D and reconstruction-based methods such as WorldTree and ViDAR.
  • It can perform visual outpainting on in-the-wild video, dynamic correction, and out-of-view dynamics inference.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)