GenRec: Knowing Where to Reconstruct and Where to Generate

Published
Source
arXiv
Paper number
947
Field
Computer Vision
arXiv ID
2608.17832

Key points

  • Training observed and unobserved pixels under a single loss makes reconstruction, which has one correct target, interfere with generation, which admits multiple plausible outputs.
  • A monocular depth estimator warps source pixels into the target view to construct an observation mask, which then controls the architecture, loss, and gradient flow.
  • A flow-matching backbone jointly reconstructs RGB and scene-coordinate maps, preserving cross-view consistency in a single sampling pass.
  • On two-view interpolation in RealEstate10K, it reached a PSNR of 19.80, well above the strongest baseline at 15.70, and achieved the best result on every metric in Mip-NeRF 360.
  • Inference takes about 11 seconds, more than two orders of magnitude faster than the strongest baseline at roughly 1,200 seconds.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)