PoEM: Predicting RL Outcomes from Existing Policies
- Published
- Source
- arXiv
- Paper number
- 1132
- Field
- Machine Learning
- arXiv ID
- 2609.30226
Key points
- They proved that if a new reward is a linear combination of existing rewards, the new log-policy can also be written as a linear combination of the existing log-policies.
- Even in non-linear cases, the empirical finding that log-policies lie in an approximately low-rank subspace became the backbone of the algorithm.
- The combination coefficients are estimated using only rewards or basis-policy outputs on samples, without running any actual reinforcement learning.
- They validated the approach on synthetic and real rewards across both text and image modalities.
- In environments like LabChin, where reward models change frequently during experiments, this can save substantial training cost.
Paper links
External research summaries. These are not HDATF publications or measured product results.