LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments

Published
Source
arXiv
Paper number
740
Field
Robotics
arXiv ID
2607.23969

Key points

  • It shifts from pixel-level video reconstruction to predicting abstract physical changes in latent space.
  • It uses a co-embedding prediction structure as a world anchor, focusing on state changes rather than visual appearance.
  • An isotropic semantic autoencoder bridges the mismatch between the predicted features and the diffusion prior.
  • The heavy prediction branch is used only during training and is removed at inference, so it runs without extra cost.
  • It achieves state-of-the-art results on RoboTwin 2.0 at 91.46 percent and on LIBERO at 97.3 percent without large-scale trajectory pretraining.
  • Isotropic regularization raises latent rank from 38.1 to 92.3 and improves success rate to 71.3 percent.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)