$τ_0$-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation
- Published
- Source
- arXiv
- Paper number
- 924
- Field
- Robotics
- arXiv ID
- 2608.16885
Key points
- The high-level policy looks at the current observation, the full instruction, previous small tasks, and execution history together to propose the next task.
- In the extra compute stage, it expands candidates over several steps, keeps only the good ones using the predicted final scene and cumulative score, and then regenerates the final task.
- Four long-horizon tasks were evaluated with 10 real robot trials per task. The mean success rate of direct execution with the same low-level policy rose from 27.5 percent to 45.0 percent in hierarchical execution without retrieval.
- With extra compute turned on, the number of successful trials across three tasks increased from 5 to 7 for making tea, from 6 to 9 for tidying books, and from 5 to 7 for cleaning a room.
- On predicting the next task for unseen book arrangements, accuracy was 50.0 percent for a single prediction, 57.5 percent for one-step candidate comparison, and 74.0 percent for multi-step extra compute. However, task-specific fine-tuning and baseline adjustment were included.
Paper links
External research summaries. These are not HDATF publications or measured product results.