PoEM: Predicting RL Outcomes from Existing Policies

Published
Source
arXiv
Paper number
1132
Field
Machine Learning
arXiv ID
2609.30226

Key points

  • They proved that if a new reward is a linear combination of existing rewards, the new log-policy can also be written as a linear combination of the existing log-policies.
  • Even in non-linear cases, the empirical finding that log-policies lie in an approximately low-rank subspace became the backbone of the algorithm.
  • The combination coefficients are estimated using only rewards or basis-policy outputs on samples, without running any actual reinforcement learning.
  • They validated the approach on synthetic and real rewards across both text and image modalities.
  • In environments like LabChin, where reward models change frequently during experiments, this can save substantial training cost.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)