InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis

Published
Source
arXiv
Paper number
806
Field
Computer Vision
arXiv ID
2608.02437

Key points

  • It built an implicit decoder that places Gaussian support points on local surfaces estimated from monocular depth and predicts Gaussian attributes from image features queried at those positions.
  • Decoupling support-point placement and attribute prediction from a fixed pixel grid produced a structure that reduces surface cracks and scattered points even under large viewpoint shifts.
  • After training only on the synthetic indoor dataset Hypersim, it outperformed existing single-image feed-forward methods on ETH3D, ScanNet++, Tanks and Temples, and DL3DV.
  • A single forward pass produces a 3D representation that supports real-time rendering, showing potential for spatial-photo exploration and augmented- and virtual-reality displays.
  • Regions unseen in a single image remain unknown, and incorrect depth estimates propagate directly to support-point placement, leaving structural distortions on reflective, transparent, or thin objects and under extreme viewpoint shifts.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)