Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling

Published
Source
arXiv
Paper number
1056
Field
Computer Vision
arXiv ID
2608.30821

Key points

  • It redistributed the requirements of a three-stage pipeline that had demanded 'precise inputs,' redesigning each stage to consume only what real captured data can reliably provide.
  • A VLM policy called GizmoAct translates object placement into 3D editor GUI operations, deciding the next adjustment and when to stop using rendered views alone.
  • It improved scene-level 3D detection mAP by 69% over Boxer, from 0.351 to 0.592.
  • It raised pose-estimation ADD-SB@0.05 on CA-1M from 57.8% to 83.4%, a gain of 25.6 percentage points, and increased the scene-reconstruction F-score from 0.794 (SAM 3D) to 0.924.
  • A single policy trained with reinforcement learning after SFT handled initialization methods with different error characteristics without retraining.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)