MilliVid: Hierarchical Latents for Long-Range Consistency in Video Generation
- Published
- Source
- arXiv
- Paper number
- 379
- Field
- Computer Vision
- arXiv ID
- 2606.09056
Key points
- It uses a hierarchical autoencoder design with different token counts per frame at different levels.
- The coarsest level, with 4 tokens per frame, captures scene structure and object permanence.
- Coarse-to-fine rollout maximizes temporal coverage within a fixed sequence length.
- Learned hierarchical tokenization provides better compression quality than simple downsampling.
- It trains the model to predict far-future frames, encouraging long-range context dependence.
- It shows a significant improvement in consistency over FramePack on long Minecraft videos.
Paper links
External research summaries. These are not HDATF publications or measured product results.