GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch

Published
Source
arXiv
Paper number
630
Field
Robotics
arXiv ID
2607.13960

Key points

  • It uses an action-centric WAM that jointly learns future visual changes and actions during training, but decodes only actions at inference without generating video.
  • It combines AC-WM and WAM pretraining and assigns visual changes and action generation to separate Transformer experts.
  • In a local RTX 4090 environment, it reduced action-inference latency to 85 milliseconds while retaining the policy-learning benefits of future visual supervision.
  • Agent-based AutoResearch automates the search over training configurations, potentially reducing the development burden of real-time closed-loop robot control.
  • The 85-millisecond figure was measured in a specific RTX 4090 deployment environment, so it does not guarantee latency or performance on other hardware or a wider variety of robot tasks.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)