Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling
- Published
- Source
- arXiv
- Paper number
- 1056
- Field
- Computer Vision
- arXiv ID
- 2608.30821
Key points
- It redistributed the requirements of a three-stage pipeline that had demanded 'precise inputs,' redesigning each stage to consume only what real captured data can reliably provide.
- A VLM policy called GizmoAct translates object placement into 3D editor GUI operations, deciding the next adjustment and when to stop using rendered views alone.
- It improved scene-level 3D detection mAP by 69% over Boxer, from 0.351 to 0.592.
- It raised pose-estimation ADD-SB@0.05 on CA-1M from 57.8% to 83.4%, a gain of 25.6 percentage points, and increased the scene-reconstruction F-score from 0.794 (SAM 3D) to 0.924.
- A single policy trained with reinforcement learning after SFT handled initialization methods with different error characteristics without retraining.
Paper links
External research summaries. These are not HDATF publications or measured product results.