SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models
- Published
- Source
- arXiv
- Paper number
- 1061
- Field
- Computer Vision
- arXiv ID
- 2609.02886
Key points
- It converted 1.43 million clips from 10 datasets, totaling at least 25TB, into a unified format containing frame alignment, camera geometry, captions, and quality metadata, enabling experiments with different training mixtures without rerunning preprocessing.
- It applied backbone-native adaptations that preserve the original representations and objectives of the structurally different Wan2.2 (5B/14B), LTX-2.5 (22B), and MiniMax-H3 (33B) backbones, producing 4 directly comparable model families.
- A simple 3-stage process of bidirectional training, autoregressive adaptation, and distribution matching distillation (DMD, a technique that compresses a large model into a smaller one by training it to imitate the large model's output distribution) achieved leading performance without special initialization.
- Despite training only on 5-second sequences, it maintained scene identity and camera responsiveness throughout continuous 60-minute rollouts without fine-tuning on long videos or using attention-sink mechanisms.
- Even when initialized with externally generated images from GPT Image 2 or Krea in styles absent from the training data, it generalized by imagining unseen regions and expanding them into explorable worlds.
- It released all data, processing pipelines, training recipes, and weights, establishing a reference point for reproducible video world-model research.
Paper links
External research summaries. These are not HDATF publications or measured product results.