Q-Learning With World Models
- Published
- Source
- arXiv
- Paper number
- 974
- Field
- Machine Learning
- arXiv ID
- 2608.17163
Key points
- It used the world model only for tree search at action-selection time, rather than as a source of training data, gaining sample efficiency without accumulating bias from imagined trajectories.
- Because the policy and Q-function are trained only on real-environment data, world-model errors do not enter the learning process.
- On the Robomimic and LIBERO robot-manipulation benchmarks, it significantly outperformed strong prior methods in both sample efficiency and final performance.
- Applying search to both training-data collection and evaluation produced the most consistent gains, and shallow search with depth 2 and a short-future weighting (lambda=0.2) was optimal.
- In image-observation (pixel-based) settings, it also generally learned faster and achieved higher success rates than the base algorithm.
Paper links
External research summaries. These are not HDATF publications or measured product results.