Geometric Action Model for Robot Policy Learning

Published
Source
arXiv
Paper number
425
Field
Robotics
arXiv ID
2606.17046

Key points

  • The GFM middle layers are split so that shallow layers contain the observation encoder and causal future predictor, while deep layers are reused as the action and geometry decoder.
  • A single autoregressive token sequence and a single forward pass generate both action tokens and future-scene tokens at the same time.
  • Compared with the existing foundation-model-scale baseline, it is 55 times faster at inference while using fewer parameters and matching or exceeding performance.
  • It is 9.7 percentage points better in the LIBERO-Plus camera perturbation setting, which reflects the direct effect of the 3D geometric prior.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)