Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents
- Published
- Source
- arXiv
- Paper number
- 684
- Field
- Robotics
- arXiv ID
- 2607.19190
Key points
- It proposes a four-agent pipeline that automatically converts real robot manipulation recordings into MuJoCo simulation episodes.
- A 31B open VLM achieves a similar conversion success rate, 48 out of 100, at only 3 percent of GPT-5.4's cost.
- It handles deformable objects such as cloth and rope, as well as humanoid motion, in the same framework, not just rigid manipulation.
- It shows that narrowing the VLM's role makes backend model replacement easier and more cost-efficient.
- The converted episode twins can be used for robot policy training and evaluation.
Paper links
External research summaries. These are not HDATF publications or measured product results.