Game2World Engine: Unlocking In-the-Wild Gameplay Videos for World Model Training
- Published
- Source
- arXiv
- Paper number
- 1010
- Field
- Computer Vision
- arXiv ID
- 2608.24680
Key points
- It automatically extracted interface fragments from real gameplay videos, collected 5,132 human-verified assets, and overlaid them on clean videos again with natural temporal behavior to synthesize paired training data.
- GameCleaner uses a multimodal language model to recognize on-screen overlays and a video diffusion transformer to restore occluded areas, so regions to erase do not need to be marked in advance.
- It achieved AAR 95.36 and background preservation 99.00 in synthetic evaluations, and AAR 80.05 and background preservation 99.80 in real-gameplay evaluations when provided with reference images.
- Training on videos with on-screen overlays removed increased the overall VideoReward score by 6.83%, opening a route to repurposing the abundance of gameplay videos on the internet for world-model training.
- However, the authors stated that reconstruction can go wrong with large opaque windows, rapidly changing overlays, or graphics entangled with the scene, and that the downstream experiment was a proxy text-to-video experiment without action conditioning.
Paper links
External research summaries. These are not HDATF publications or measured product results.