HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL

Published
Source
arXiv
Paper number
937
Field
Robotics
arXiv ID
2608.16837

Key points

  • HAF-VLA activates head and locomotion, the waist, and bimanual manipulation sequentially across three stages. Stage indicators and the previous stage's KV cache connect earlier body movements with subsequent manipulation actions.
  • HAF-Steer compresses the VLA's temporal noise with a discrete cosine transform and learns only the first eight coefficients. It begins with behavior cloning, then applies SAC reinforcement learning that mixes offline data with real interactions while keeping the VLA backbone frozen.
  • The researchers tested seven tasks on TienKung 2.0 and 3.0, including loading laundry, carrying a basket, and throwing a ball. They collected 120 teleoperated demonstrations for each task and ran each method on each task ten times.
  • Across the seven tasks, HAF-VLA achieved an average progress score of 70.5%, compared with 53.3% for the strongest baseline, π0.5. It recorded the highest or joint-highest score on every task.
  • HAF-Steer improved HAF-VLA's success rate across all four conditions covering the original and new locations for two tasks. For example, toy storage at the new location improved from eight successes in ten trials to ten in ten. Its limitations include additional computational latency and dependence on the action range of the base VLA.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)