Programmable World Model

Published
Source
arXiv
Paper number
1078
Field
Computer Vision
arXiv ID
2609.10540

Key points

  • Separating a state-managing engine from a video-generating renderer structurally solves video world models' inability to remember.
  • A coding agent turns natural-language instructions into executable programs, letting users directly program game rules.
  • A 3D oriented-bounding-box intermediate representation keeps programs easy to manipulate while precisely guiding video generation.
  • On the new CombatStateBench it reaches 94% character-count consistency and 98% state fidelity, leading prior models by up to 62 percentage points.
  • It also stays best on all four VBench video-quality metrics, showing structured control does not hurt visual quality.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)