Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

Published
Source
arXiv
Paper number
603
Field
Robotics
arXiv ID
2607.11643

Key points

  • It trains five generative tasks, image generation, editing, embodied scene generation, transfer, and video generation, in a single 38B-parameter autoregressive framework.
  • A structured control format separates workspace, background, foreground, target, and lighting as independent control dimensions to preserve multi-view consistency.
  • GPT-Image-2.0 ranks first in human evaluation for embodied scene generation and transfer, and first in World Arena embodied video generation.
  • Synthetic data generated by the model improves the OOD success rate of the π0.5 robot policy from 36.9% to 63.2%, a 26.3-point gain.
  • Sequential world modeling turns static scene generation into a scalable trajectory generation engine.
  • It releases code and checkpoints so the model can serve as an embodied-intelligence data engine.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)