Code World Model: Coding Agent as World Brain
- Published
- Source
- arXiv
- Paper number
- 1017
- Field
- Computer Vision
- arXiv ID
- 2608.25927
Key points
- It diagnosed existing world models trained only on video as imitating resulting scenes without learning the rules that explain why things happen.
- It divided responsibilities into a dual structure: a coding agent, the brain, makes infrequent complex decisions, while code rapidly handles frequently repeated state updates.
- It extracted 9,420 clips from 157 gameplay takes, approximately 5.6 hours, and fine-tuned MiniMax-H3 with LoRA using approximately 596 million trainable parameters.
- It reported that the proxy-conditioning resolution is approximately 1/16 of the output resolution, adding almost no inference overhead.
- It acknowledged limitations: the training scale is small, generation is not real-time, and it remains difficult for the coding agent to create complex game mechanics from scratch.
Paper links
External research summaries. These are not HDATF publications or measured product results.