Closing the Loop in Humanoid VLA: Persistent 3D Object Tokens for Verifiable Loco-Manipulation

Published
Source
arXiv
Paper number
678
Field
Robotics
arXiv ID
2607.18016

Key points

  • It points out that robots fail on long tasks not because they cannot understand instructions, but because the object state used for action differs from the object state used for verification.
  • POT binds objects to role slots such as target, destination, support surface, and handoff partner, and refreshes the 3D position grounding after each action chunk.
  • Object records are turned into fixed-format tokens containing 33 features per slot across 8 slots and inserted into the input sequence of the whole-body action model.
  • The same record is used to verify grasp, placement, and release through geometry-based checks, and if it fails, the system triggers re-observation, retry, and replanning.
  • On eight task groups with a real Unitree G1, it improves the same-condition GR00T-N1.7 baseline from 39/80 to 71/80, and cup stacking jumps from 1/10 to 8/10.
  • In ablation, tokens alone improve from 15/40 to 31/40, verification alone gives 22/40, and using both gives 34/40, showing that token conditioning is the main driver of performance.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)