Code World Models for General Game Playing
- Published
- Source
- arXiv
- Paper number
- 459
- Field
- Agents / World Models
- arXiv ID
- 2510.04542
Key points
- The core method uses the LLM not as a policy that directly chooses moves, but as an inductive engine that translates rules and trajectory data into Python code. The generated Code World Model (CWM) is an executable simulation engine that includes state-transition, legal-move enumeration, and terminal-check functions.
- On top of the CWM, it places MCTS for perfect-information games or ISMCTS for imperfect-information games to perform deep search-based planning. In addition, the LLM is asked to generate a heuristic value function and a hidden-state inference function as code.
- For imperfect-information games, it proposes a code-based autoencoder paradigm. The inference function serves as the encoder from observation to hidden state, the CWM serves as the decoder from hidden state to observation, and the game rules act as a structural regularizer.
- It was evaluated on 10 games, five with perfect information and five with imperfect information, including four OOD games newly created for the paper. CWM-MCTS and CWM-ISMCTS performed nearly as well as players based on ground-truth models, and they beat Gemini 2.5 Pro in 9 of the 10 games or tied with it.
- In perfect-information games, CWM transition accuracy reached 1.0 in most cases, while in imperfect-information games hidden-state inference remained difficult and accuracy dropped to around 52% on the hard game Gin Rummy.
- The limitations are that learning remains difficult in imperfect-information games with complex rules such as Gin Rummy, and extending the approach to open-world or vision-based games remains future work.
Paper links
External research summaries. These are not HDATF publications or measured product results.