From Foundation to Application: Improving VLA Models in Practice

Published
Source
arXiv
Paper number
574
Field
Robotics
arXiv ID
2607.06403

Key points

  • It builds about 60,000 hours of pretraining data, made up of 50,000 hours of robot trajectories and 10,000 hours of egocentric video, across 20 robot configurations.
  • It supports control of the full body, extending beyond the arms to the head, waist, mobile base, and active hands.
  • It introduces dual-query distillation for predictive dynamics modeling, using a video representation model as the semantic prior and a depth estimation model as the geometric cue.
  • On the GM-100 dual-arm benchmark, it achieves a 55.0% average success rate across four tasks using MeanStd normalization and L2 loss, with the best task reaching 76.8%.
  • The relative action target, relQpos, compresses the action scale to 0.34x compared with the absolute target and improves training stability.
  • It applies token-level adaptive expert routing with a MoE-based action expert structure.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)