Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning
- Published
- Source
- arXiv
- Paper number
- 864
- Field
- Computer Vision
- arXiv ID
- 2608.09926
Key points
- We propose LDR, the first video world model to explicitly infer the laws of motion in latent space.
- The architecture enforces learning of physical laws by numerically integrating the lower-order dynamics and having the model regress only the third- and higher-order residual that drives the rollout.
- It operates with 26x fewer parameters and 143x faster speed than video diffusion models.
- Its out-of-distribution error gap is more than 20x smaller, showing strong extrapolation ability.
- It accurately generalizes learned dynamics even when color, shape, and orientation change.
Paper links
External research summaries. These are not HDATF publications or measured product results.