Geometric Action Model for Robot Policy Learning
- Published
- Source
- arXiv
- Paper number
- 425
- Field
- Robotics
- arXiv ID
- 2606.17046
Key points
- The GFM middle layers are split so that shallow layers contain the observation encoder and causal future predictor, while deep layers are reused as the action and geometry decoder.
- A single autoregressive token sequence and a single forward pass generate both action tokens and future-scene tokens at the same time.
- Compared with the existing foundation-model-scale baseline, it is 55 times faster at inference while using fewer parameters and matching or exceeding performance.
- It is 9.7 percentage points better in the LIBERO-Plus camera perturbation setting, which reflects the direct effect of the 3D geometric prior.
Paper links
External research summaries. These are not HDATF publications or measured product results.