H3-World: Turning Language Understanding into World Control
- Published
- Source
- arXiv
- Paper number
- 1057
- Field
- Computer Vision
- arXiv ID
- 2609.01560
Key points
- Rather than attaching a dedicated control module to the 33B video-generation model MiniMax-H3, it directly reused the text pathway the model already understands as a control interface.
- It achieved precise character and camera control using just 8,000 gameplay samples, 10,000 LoRA training steps, and 0.199% trainable parameters.
- Temporal attention routing restricts each command to its own time interval, reducing command leakage into other intervals.
- In camera-direction-change experiments, the global-prompt approach achieved horizontal flow of only +0.0/-17.3, whereas H3-World accurately followed both directions with +52.7/-106.0.
- The same interface worked unchanged for action combinations never seen during training and for scenes outside games, including first-person views, interiors, and fantasy settings.
Paper links
External research summaries. These are not HDATF publications or measured product results.