Reversal Q-Learning
- Published
- Source
- arXiv
- Paper number
- 453
- Field
- Machine Learning
- arXiv ID
- 2606.17551
Key points
- It applies an extended MDP framework to offline RL and generates virtual on-policy trajectories through a reverse flow.
- Multi-step returns reduce the effective horizon from T×F to T, solving the curse of horizon.
- It uses first-order gradients of the value function without BPTT, outperforming regression-based methods such as FAWAC and IFQL.
- It achieves the best average performance among 18 baselines on 50 robot tasks in OGBench.
- It is especially strong on long-horizon tasks such as antmaze-giant and humanoidmaze-large.
Paper links
External research summaries. These are not HDATF publications or measured product results.