Sekai2: From World Exploration to Interactive World Modeling

Published
Source
arXiv
Paper number
863
Field
Computer Vision
arXiv ID
2608.09449

Key points

  • It provides 2,826 hours of real-world video, 128,892 clips, and camera trajectories with hierarchical annotations across 113 countries.
  • It creates 649,597 time-level segments that separate subject motion, environmental change, static scenes, and camera behavior.
  • It offers 982 loop videos revisiting 360-degree panoramas, which provide supervision for long-range spatial consistency.
  • It cross-validates that the captions generated independently of camera annotations are aligned with actual camera motion.
  • It provides a scalable data foundation for video generation, camera control, and pretraining interactive world models.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)