MilliVid: Hierarchical Latents for Long-Range Consistency in Video Generation

Published
Source
arXiv
Paper number
379
Field
Computer Vision
arXiv ID
2606.09056

Key points

  • It uses a hierarchical autoencoder design with different token counts per frame at different levels.
  • The coarsest level, with 4 tokens per frame, captures scene structure and object permanence.
  • Coarse-to-fine rollout maximizes temporal coverage within a fixed sequence length.
  • Learned hierarchical tokenization provides better compression quality than simple downsampling.
  • It trains the model to predict far-future frames, encouraging long-range context dependence.
  • It shows a significant improvement in consistency over FramePack on long Minecraft videos.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)