Game2World Engine: Unlocking In-the-Wild Gameplay Videos for World Model Training

Published
Source
arXiv
Paper number
1010
Field
Computer Vision
arXiv ID
2608.24680

Key points

  • It automatically extracted interface fragments from real gameplay videos, collected 5,132 human-verified assets, and overlaid them on clean videos again with natural temporal behavior to synthesize paired training data.
  • GameCleaner uses a multimodal language model to recognize on-screen overlays and a video diffusion transformer to restore occluded areas, so regions to erase do not need to be marked in advance.
  • It achieved AAR 95.36 and background preservation 99.00 in synthetic evaluations, and AAR 80.05 and background preservation 99.80 in real-gameplay evaluations when provided with reference images.
  • Training on videos with on-screen overlays removed increased the overall VideoReward score by 6.83%, opening a route to repurposing the abundance of gameplay videos on the internet for world-model training.
  • However, the authors stated that reconstruction can go wrong with large opaque windows, rapidly changing overlays, or graphics entangled with the scene, and that the downstream experiment was a proxy text-to-video experiment without action conditioning.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)