SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models

Published
Source
arXiv
Paper number
1061
Field
Computer Vision
arXiv ID
2609.02886

Key points

  • It converted 1.43 million clips from 10 datasets, totaling at least 25TB, into a unified format containing frame alignment, camera geometry, captions, and quality metadata, enabling experiments with different training mixtures without rerunning preprocessing.
  • It applied backbone-native adaptations that preserve the original representations and objectives of the structurally different Wan2.2 (5B/14B), LTX-2.5 (22B), and MiniMax-H3 (33B) backbones, producing 4 directly comparable model families.
  • A simple 3-stage process of bidirectional training, autoregressive adaptation, and distribution matching distillation (DMD, a technique that compresses a large model into a smaller one by training it to imitate the large model's output distribution) achieved leading performance without special initialization.
  • Despite training only on 5-second sequences, it maintained scene identity and camera responsiveness throughout continuous 60-minute rollouts without fine-tuning on long videos or using attention-sink mechanisms.
  • Even when initialized with externally generated images from GPT Image 2 or Krea in styles absent from the training data, it generalized by imagining unseen regions and expanding them into explorable worlds.
  • It released all data, processing pipelines, training recipes, and weights, establishing a reference point for reproducible video world-model research.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)