$τ_0$-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation

Published
Source
arXiv
Paper number
924
Field
Robotics
arXiv ID
2608.16885

Key points

  • The high-level policy looks at the current observation, the full instruction, previous small tasks, and execution history together to propose the next task.
  • In the extra compute stage, it expands candidates over several steps, keeps only the good ones using the predicted final scene and cumulative score, and then regenerates the final task.
  • Four long-horizon tasks were evaluated with 10 real robot trials per task. The mean success rate of direct execution with the same low-level policy rose from 27.5 percent to 45.0 percent in hierarchical execution without retrieval.
  • With extra compute turned on, the number of successful trials across three tasks increased from 5 to 7 for making tea, from 6 to 9 for tidying books, and from 5 to 7 for cleaning a room.
  • On predicting the next task for unseen book arrangements, accuracy was 50.0 percent for a single prediction, 57.5 percent for one-step candidate comparison, and 74.0 percent for multi-step extra compute. However, task-specific fine-tuning and baseline adjustment were included.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)