Q-Learning With World Models

Published
Source
arXiv
Paper number
974
Field
Machine Learning
arXiv ID
2608.17163

Key points

  • It used the world model only for tree search at action-selection time, rather than as a source of training data, gaining sample efficiency without accumulating bias from imagined trajectories.
  • Because the policy and Q-function are trained only on real-environment data, world-model errors do not enter the learning process.
  • On the Robomimic and LIBERO robot-manipulation benchmarks, it significantly outperformed strong prior methods in both sample efficiency and final performance.
  • Applying search to both training-data collection and evaluation produced the most consistent gains, and shallow search with depth 2 and a short-future weighting (lambda=0.2) was optimal.
  • In image-observation (pixel-based) settings, it also generally learned faster and achieved higher success rates than the base algorithm.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)