FixAnything: 3D-Consistent Rendering Refinement via Video Generative Priors

Published
Source
arXiv
Paper number
997
Field
Computer Vision
arXiv ID
2608.23549

Key points

  • It proposed a general-purpose framework that cleans up rendering defects across 4 representations, 3DGS, NeRF, meshes, and point clouds, using a single model.
  • However, the model must invent regions not visible in the input views, and the paper's uncertainty analysis also showed a large quality gap: the 25% of pixels with high confidence had a PSNR of 25.7dB, while the 25% with low confidence had 14.4dB.
  • Mask conditioning that identifies clean pixels distinguishes frames to trust from frames to repair, preventing hallucinations; removing the mask caused a 1.3dB drop.
  • Applying Flow-DPO with camera-pose reconstruction accuracy as the reward improved 3D consistency, measured by AUC@5°, by 7.2%.
  • It matches or exceeds expert pipelines in quality, with an architecture that only requires swapping in a stronger video model when one becomes available.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)