Thinking in Blender: Staged Executable Inverse Graphics with Vision-Language Models

Published
Source
arXiv
Paper number
285
Field
Computer Vision
arXiv ID
2606.02580

Key points

  • The paper introduces Staged Executable Inverse Graphics, or SEIG, an agent framework for reconstructing 3D scenes from a single image by progressively refining scene factors, including geometry, materials, composition, and lighting, directly in executable Blender code space.
  • The experiments show that staged reconstruction substantially improves reconstruction fidelity and highlights the importance of task decomposition in executable inverse graphics with a general-purpose VLM.
  • It finally demonstrates a variety of downstream applications enabled by the reconstructed and editable Blender scenes.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)