From Foundation to Application: Improving VLA Models in Practice
- Published
- Source
- arXiv
- Paper number
- 574
- Field
- Robotics
- arXiv ID
- 2607.06403
Key points
- It builds about 60,000 hours of pretraining data, made up of 50,000 hours of robot trajectories and 10,000 hours of egocentric video, across 20 robot configurations.
- It supports control of the full body, extending beyond the arms to the head, waist, mobile base, and active hands.
- It introduces dual-query distillation for predictive dynamics modeling, using a video representation model as the semantic prior and a depth estimation model as the geometric cue.
- On the GM-100 dual-arm benchmark, it achieves a 55.0% average success rate across four tasks using MeanStd normalization and L2 loss, with the best task reaching 76.8%.
- The relative action target, relQpos, compresses the action scale to 0.34x compared with the absolute target and improves training stability.
- It applies token-level adaptive expert routing with a MoE-based action expert structure.
Paper links
External research summaries. These are not HDATF publications or measured product results.