LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments
- Published
- Source
- arXiv
- Paper number
- 740
- Field
- Robotics
- arXiv ID
- 2607.23969
Key points
- It shifts from pixel-level video reconstruction to predicting abstract physical changes in latent space.
- It uses a co-embedding prediction structure as a world anchor, focusing on state changes rather than visual appearance.
- An isotropic semantic autoencoder bridges the mismatch between the predicted features and the diffusion prior.
- The heavy prediction branch is used only during training and is removed at inference, so it runs without extra cost.
- It achieves state-of-the-art results on RoboTwin 2.0 at 91.46 percent and on LIBERO at 97.3 percent without large-scale trajectory pretraining.
- Isotropic regularization raises latent rank from 38.1 to 92.3 and improves success rate to 71.3 percent.
Paper links
External research summaries. These are not HDATF publications or measured product results.