Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack

Published
Source
arXiv
Paper number
415
Field
Robotics
arXiv ID
2606.14409

Key points

  • It collects 10,000 hours of sub-millimeter-precision egocentric UMI data in-house and reuses it for both pretraining and post-training.
  • It combines the Hy-Embodied-0.5 MoT backbone with a flow-matching action expert and a compact memory encoder, using a delta-chunk representation to achieve kinematics independence.
  • FlowPRO is an offline RL post-training method that learns directly from success-failure pairs without a critic or reward model and reaches near-ceiling success rates.
  • It uses two SFT tracks, Track A for adaptation to a target robot and Track B for UMI-only cross-embodiment transfer, to deploy across diverse robot platforms.
  • It implements high-frequency closed-loop control with asynchronous inference and cubic Bézier trajectory smoothing, and it plans to release a 2,000-hour UMI data subset.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)