Code World Model: Coding Agent as World Brain

Published
Source
arXiv
Paper number
1017
Field
Computer Vision
arXiv ID
2608.25927

Key points

  • It diagnosed existing world models trained only on video as imitating resulting scenes without learning the rules that explain why things happen.
  • It divided responsibilities into a dual structure: a coding agent, the brain, makes infrequent complex decisions, while code rapidly handles frequently repeated state updates.
  • It extracted 9,420 clips from 157 gameplay takes, approximately 5.6 hours, and fine-tuned MiniMax-H3 with LoRA using approximately 596 million trainable parameters.
  • It reported that the proxy-conditioning resolution is approximately 1/16 of the output resolution, adding almost no inference overhead.
  • It acknowledged limitations: the training scale is small, generation is not real-time, and it remains difficult for the coding agent to create complex game mechanics from scratch.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)