Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
- Published
- Source
- arXiv
- Paper number
- 581
- Field
- Computer Vision
- arXiv ID
- 2607.07675
Key points
- The MoE, or sparse expert, architecture improves the trade-off between inference efficiency and model capacity relative to dense models.
- The data profiling engine jointly analyzes, filters, and rebalances internet video and robot manipulation, navigation, and egocentric data.
- A multidimensional reward system aligns not only aesthetic quality but also physical rationality and task-completion signals during training.
- A single-stream Diffusion Transformer with Multi-Modal 3D RoPE unifies T2I, T2V, and I2V within one framework.
- QK-Norm and AdaLN-Single modulation stabilize deep transformers, and a cascaded refiner improves quality.
- It is the first large-scale open-source MoE video model for embodied intelligence, and the code and checkpoints are released.
Paper links
External research summaries. These are not HDATF publications or measured product results.