Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization

Published
Source
arXiv
Paper number
255
Field
Machine Learning
arXiv ID
2605.26282

Key points

  • Identify the structural mismatch between exploration and value learning as the key bottleneck in world-model RL.
  • Recast policy optimization as a diffusion process on the latent world model to unify exploration and policy optimization.
  • Anchor the policy to the behavior distribution with a dataset-induced implicit energy function, namely a KL trust region.
  • Correct the score field with cumulative reward from imagined trajectories so it converges toward the optimal Gibbs policy.
  • Show consistent gains across multi-task offline pretraining, online, and offline-to-online settings.
  • Demonstrate scalability through monotonic performance improvements as model capacity increases.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)