Flex-$π$: A Multi-Stream World-Action Model with Compute Flexibility

Published
Source
arXiv
Paper number
881
Field
Robotics
arXiv ID
2608.10860

Key points

  • It discovers a free lunch in which a video-generation model's VAE can encode not only RGB but also 3D point maps with almost no loss.
  • It builds a multi-stream world-action model that predicts RGB, 3D depth, and DINO object features together in a shared latent space.
  • During training it randomly drops modalities, or stream dropout, so that one checkpoint works with any combination at inference time.
  • On robot tasks in RoboTwin, with only 50 demos, it is 1.9x better than existing WAMs, and on a real dual-arm robot it achieves 2x to 7x higher success.
  • In action-only mode it is faster than π0.5 and still achieves higher success, at about 60 ms.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)