Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild
- Published
- Source
- arXiv
- Paper number
- 484
- Field
- Computer Vision
- arXiv ID
- 2606.23688
Key points
- We apply causal latent conditioning to a monocular 3D reconstruction model to ensure temporal consistency across frames.
- We convert temporally consistent initialization into a deformable 3D Gaussian Splatting representation.
- We use occlusion-aware rendering supervision, which localizes occluded regions with depth cues and handles visible and invisible regions separately.
- We apply a view-conditioned diffusion prior to the unseen regions to complete plausible appearances.
- On the Consistent4D synthetic benchmark, we achieve SOTA with LPIPS 0.116 and FVD 592.44.
- On DAVIS in-the-wild videos, we demonstrate superior motion accuracy over prior methods, with CLIP 0.715 and EPE 0.161.
Paper links
External research summaries. These are not HDATF publications or measured product results.