H-JEPA: End-to-End Learning of Hierarchical World Models for Visual Planning

Published
Source
arXiv
Paper number
1169
Field
Machine Learning
arXiv ID
2610.06805

Key points

  • The world model is stacked into three levels instead of one, with each level predicting the future at a different temporal stride (5/10/20 steps).
  • Higher levels discard fast-changing detail and keep only slowly varying, important state, learning representations that match abstract goals well.
  • On the visual maze-navigation task, the three-level structure raised the success rate from 18% to 73% over the flat single-level structure while using less planner compute.
  • The hierarchy emerged purely from prediction learning, with no human-provided reward and no pixel reconstruction.
  • With an inverse-dynamics auxiliary loss, hierarchy also improved planning fidelity on the real-robot video dataset DROID.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)