Qwen-AgentWorld: Language World Models for General Agents
- Published
- Source
- arXiv
- Paper number
- 480
- Field
- World Models / Agents
- arXiv ID
- 2606.24597
Key points
- It presents language world models, Qwen-AgentWorld-35B-A3B and Qwen-AgentWorld-397B-A17B, across seven simulated domains.
- It uses a three-stage training process, CPT for world-model capability injection, SFT for activating next-state prediction, and RL for improving simulation accuracy.
- It introduces AgentWorldBench and shows superior performance across five frontier models and nine benchmarks.
- In Paradigm A, the separated approach, thousands of environment simulations for agent RL outperform training in the real environment.
- In Paradigm B, the integrated approach, world-model training is used as warmup and improves performance on seven agent benchmarks.
- It trains on more than 10M environment interaction trajectories and supports reasoning based on long chains of thought.
Paper links
External research summaries. These are not HDATF publications or measured product results.