Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model
- Published
- Source
- arXiv
- Paper number
- 603
- Field
- Robotics
- arXiv ID
- 2607.11643
Key points
- It trains five generative tasks, image generation, editing, embodied scene generation, transfer, and video generation, in a single 38B-parameter autoregressive framework.
- A structured control format separates workspace, background, foreground, target, and lighting as independent control dimensions to preserve multi-view consistency.
- GPT-Image-2.0 ranks first in human evaluation for embodied scene generation and transfer, and first in World Arena embodied video generation.
- Synthetic data generated by the model improves the OOD success rate of the π0.5 robot policy from 36.9% to 63.2%, a 26.3-point gain.
- Sequential world modeling turns static scene generation into a scalable trajectory generation engine.
- It releases code and checkpoints so the model can serve as an embodied-intelligence data engine.
Paper links
External research summaries. These are not HDATF publications or measured product results.