Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization
- Published
- Source
- arXiv
- Paper number
- 255
- Field
- Machine Learning
- arXiv ID
- 2605.26282
Key points
- Identify the structural mismatch between exploration and value learning as the key bottleneck in world-model RL.
- Recast policy optimization as a diffusion process on the latent world model to unify exploration and policy optimization.
- Anchor the policy to the behavior distribution with a dataset-induced implicit energy function, namely a KL trust region.
- Correct the score field with cumulative reward from imagined trajectories so it converges toward the optimal Gibbs policy.
- Show consistent gains across multi-task offline pretraining, online, and offline-to-online settings.
- Demonstrate scalability through monotonic performance improvements as model capacity increases.
Paper links
External research summaries. These are not HDATF publications or measured product results.