Latent Spatial Memory for Video World Models

Published
Source
arXiv
Paper number
371
Field
Computer Vision
arXiv ID
2606.09828

Key points

  • We build 3D spatial memory directly in latent space, eliminating the cost of round-tripping through pixel space.
  • Depth-guided back-projection loads latent tokens into 3D world space.
  • Occlusion-aware latent-resolution projection is used to obtain the target view.
  • We exclude dynamic objects and the sky from memory updates to improve long-term stability.
  • Compared with an RGB cache, the system runs 10.57 times faster end to end and uses 55 times less GPU memory.
  • Two-stage training enables stable convergence of the side branch and LoRA.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)