Guava: An Effective and Universal Harness for Embodied Manipulation

Published
Source
arXiv
Paper number
452
Field
Robotics
arXiv ID
2606.18363

Key points

  • ReAct-style closed-loop execution is decisively better than one-shot planning on long-horizon tasks.
  • A semantic action space, such as grasp, align, and move, reduces the reasoning burden on the VLM compared with low-level joint control.
  • Training on fewer than 2,000 simulation trajectories and then transferring zero-shot to the real world succeeds on three backbones, with 86 percent on in-distribution tasks and 92 percent on out-of-distribution tasks.
  • RL post-training raises shell game success from 6.7 percent to 60 percent and place-all-red-objects success from 0 percent to 93.3 percent.
  • The system exhibits error awareness and recovery behavior even in joint-limit or unreachable situations that were absent from training.
  • It achieves a 75.6 percent overall success rate, outperforming GPT-5.4 at 70.2 percent and CaP-Agent0 at 62.7 percent.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)