Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

Published
Source
arXiv
Paper number
581
Field
Computer Vision
arXiv ID
2607.07675

Key points

  • The MoE, or sparse expert, architecture improves the trade-off between inference efficiency and model capacity relative to dense models.
  • The data profiling engine jointly analyzes, filters, and rebalances internet video and robot manipulation, navigation, and egocentric data.
  • A multidimensional reward system aligns not only aesthetic quality but also physical rationality and task-completion signals during training.
  • A single-stream Diffusion Transformer with Multi-Modal 3D RoPE unifies T2I, T2V, and I2V within one framework.
  • QK-Norm and AdaLN-Single modulation stabilize deep transformers, and a cascaded refiner improves quality.
  • It is the first large-scale open-source MoE video model for embodied intelligence, and the code and checkpoints are released.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)