Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Published
Source
arXiv
Paper number
439
Field
Robotics
arXiv ID
2606.17846

Key points

  • It proposes a three-dimensional alignment framework for representation, motion, and action, using canonical state, camera-frame delta pose, and in-context adaptation.
  • It builds a human-to-robot synthesis pipeline that converts human egocentric video into trajectories for 15 robot platforms.
  • It builds a pretraining corpus of about 38,100 hours of manipulation data using only open-source data, with no proprietary data collection.
  • In out-of-distribution evaluation on LIBERO-Plus, RoboCasa365, and EBench, it outperforms π0.5 and GR00T-N1.7 across all axes.
  • It ranks first on the RoboChallenge Table30-v1 generalist track and achieves a 20 percent relative improvement.
  • It has been validated on four real robot platforms: AgileX ALOHA, Franka, UR, and ARX.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)