Qwen-AgentWorld: Language World Models for General Agents

Published
Source
arXiv
Paper number
480
Field
World Models / Agents
arXiv ID
2606.24597

Key points

  • It presents language world models, Qwen-AgentWorld-35B-A3B and Qwen-AgentWorld-397B-A17B, across seven simulated domains.
  • It uses a three-stage training process, CPT for world-model capability injection, SFT for activating next-state prediction, and RL for improving simulation accuracy.
  • It introduces AgentWorldBench and shows superior performance across five frontier models and nine benchmarks.
  • In Paradigm A, the separated approach, thousands of environment simulations for agent RL outperform training in the real environment.
  • In Paradigm B, the integrated approach, world-model training is used as warmup and improves performance on seven agent benchmarks.
  • It trains on more than 10M environment interaction trajectories and supports reasoning based on long chains of thought.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)