Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

Published
Source
arXiv
Paper number
660
Field
Robotics
arXiv ID
2607.15330

Key points

  • An automatic VLM-based labeling pipeline is applied to more than 100,000 hours of real UMI manipulation trajectories.
  • The pretraining stage shows a clear scaling law, with performance improving consistently as both data and model size increase.
  • On RoboCasa365, it reaches 57.6 percent, which beats the previous best of 46.6 percent.
  • On a real robot, it learns new tasks from less than 10 hours of data and reaches an average success rate of 75 percent.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)