Closing the Loop in Humanoid VLA: Persistent 3D Object Tokens for Verifiable Loco-Manipulation
- Published
- Source
- arXiv
- Paper number
- 678
- Field
- Robotics
- arXiv ID
- 2607.18016
Key points
- It points out that robots fail on long tasks not because they cannot understand instructions, but because the object state used for action differs from the object state used for verification.
- POT binds objects to role slots such as target, destination, support surface, and handoff partner, and refreshes the 3D position grounding after each action chunk.
- Object records are turned into fixed-format tokens containing 33 features per slot across 8 slots and inserted into the input sequence of the whole-body action model.
- The same record is used to verify grasp, placement, and release through geometry-based checks, and if it fails, the system triggers re-observation, retry, and replanning.
- On eight task groups with a real Unitree G1, it improves the same-condition GR00T-N1.7 baseline from 39/80 to 71/80, and cup stacking jumps from 1/10 to 8/10.
- In ablation, tokens alone improve from 15/40 to 31/40, verification alone gives 22/40, and using both gives 34/40, showing that token conditioning is the main driver of performance.
Paper links
External research summaries. These are not HDATF publications or measured product results.