Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning

Published
Source
arXiv
Paper number
864
Field
Computer Vision
arXiv ID
2608.09926

Key points

  • We propose LDR, the first video world model to explicitly infer the laws of motion in latent space.
  • The architecture enforces learning of physical laws by numerically integrating the lower-order dynamics and having the model regress only the third- and higher-order residual that drives the rollout.
  • It operates with 26x fewer parameters and 143x faster speed than video diffusion models.
  • Its out-of-distribution error gap is more than 20x smaller, showing strong extrapolation ability.
  • It accurately generalizes learned dynamics even when color, shape, and orientation change.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)