H3-World: Turning Language Understanding into World Control

Published
Source
arXiv
Paper number
1057
Field
Computer Vision
arXiv ID
2609.01560

Key points

  • Rather than attaching a dedicated control module to the 33B video-generation model MiniMax-H3, it directly reused the text pathway the model already understands as a control interface.
  • It achieved precise character and camera control using just 8,000 gameplay samples, 10,000 LoRA training steps, and 0.199% trainable parameters.
  • Temporal attention routing restricts each command to its own time interval, reducing command leakage into other intervals.
  • In camera-direction-change experiments, the global-prompt approach achieved horizontal flow of only +0.0/-17.3, whereas H3-World accurately followed both directions with +52.7/-106.0.
  • The same interface worked unchanged for action combinations never seen during training and for scenes outside games, including first-person views, interiors, and fantasy settings.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)