Flex-$π$: A Multi-Stream World-Action Model with Compute Flexibility
- Published
- Source
- arXiv
- Paper number
- 881
- Field
- Robotics
- arXiv ID
- 2608.10860
Key points
- It discovers a free lunch in which a video-generation model's VAE can encode not only RGB but also 3D point maps with almost no loss.
- It builds a multi-stream world-action model that predicts RGB, 3D depth, and DINO object features together in a shared latent space.
- During training it randomly drops modalities, or stream dropout, so that one checkpoint works with any combination at inference time.
- On robot tasks in RoboTwin, with only 50 demos, it is 1.9x better than existing WAMs, and on a real dual-arm robot it achieves 2x to 7x higher success.
- In action-only mode it is faster than π0.5 and still achieves higher success, at about 60 ms.
Paper links
External research summaries. These are not HDATF publications or measured product results.