Reversal Q-Learning

Published
Source
arXiv
Paper number
453
Field
Machine Learning
arXiv ID
2606.17551

Key points

  • It applies an extended MDP framework to offline RL and generates virtual on-policy trajectories through a reverse flow.
  • Multi-step returns reduce the effective horizon from T×F to T, solving the curse of horizon.
  • It uses first-order gradients of the value function without BPTT, outperforming regression-based methods such as FAWAC and IFQL.
  • It achieves the best average performance among 18 baselines on 50 robot tasks in OGBench.
  • It is especially strong on long-horizon tasks such as antmaze-giant and humanoidmaze-large.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)