Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild

Published
Source
arXiv
Paper number
484
Field
Computer Vision
arXiv ID
2606.23688

Key points

  • We apply causal latent conditioning to a monocular 3D reconstruction model to ensure temporal consistency across frames.
  • We convert temporally consistent initialization into a deformable 3D Gaussian Splatting representation.
  • We use occlusion-aware rendering supervision, which localizes occluded regions with depth cues and handles visible and invisible regions separately.
  • We apply a view-conditioned diffusion prior to the unseen regions to complete plausible appearances.
  • On the Consistent4D synthetic benchmark, we achieve SOTA with LPIPS 0.116 and FVD 592.44.
  • On DAVIS in-the-wild videos, we demonstrate superior motion accuracy over prior methods, with CLIP 0.715 and EPE 0.161.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)