Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents

Published
Source
arXiv
Paper number
684
Field
Robotics
arXiv ID
2607.19190

Key points

  • It proposes a four-agent pipeline that automatically converts real robot manipulation recordings into MuJoCo simulation episodes.
  • A 31B open VLM achieves a similar conversion success rate, 48 out of 100, at only 3 percent of GPT-5.4's cost.
  • It handles deformable objects such as cloth and rope, as well as humanoid motion, in the same framework, not just rigid manipulation.
  • It shows that narrowing the VLM's role makes backend model replacement easier and more cost-efficient.
  • The converted episode twins can be used for robot policy training and evaluation.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)