Latent Spatial Memory for Video World Models
- Published
- Source
- arXiv
- Paper number
- 371
- Field
- Computer Vision
- arXiv ID
- 2606.09828
Key points
- We build 3D spatial memory directly in latent space, eliminating the cost of round-tripping through pixel space.
- Depth-guided back-projection loads latent tokens into 3D world space.
- Occlusion-aware latent-resolution projection is used to obtain the target view.
- We exclude dynamic objects and the sky from memory updates to improve long-term stability.
- Compared with an RGB cache, the system runs 10.57 times faster end to end and uses 55 times less GPU memory.
- Two-stage training enables stable convergence of the side branch and LoRA.
Paper links
External research summaries. These are not HDATF publications or measured product results.