H-JEPA: End-to-End Learning of Hierarchical World Models for Visual Planning
- Published
- Source
- arXiv
- Paper number
- 1169
- Field
- Machine Learning
- arXiv ID
- 2610.06805
Key points
- The world model is stacked into three levels instead of one, with each level predicting the future at a different temporal stride (5/10/20 steps).
- Higher levels discard fast-changing detail and keep only slowly varying, important state, learning representations that match abstract goals well.
- On the visual maze-navigation task, the three-level structure raised the success rate from 18% to 73% over the flat single-level structure while using less planner compute.
- The hierarchy emerged purely from prediction learning, with no human-provided reward and no pixel reconstruction.
- With an inverse-dynamics auxiliary loss, hierarchy also improved planning fidelity on the real-robot video dataset DROID.
Paper links
External research summaries. These are not HDATF publications or measured product results.