Regret Minimization with Adaptive Opponents in Repeated Games
- Published
- Source
- arXiv
- Paper number
- 329
- Field
- Machine Learning
- arXiv ID
- 2606.06486
Key points
- It introduces RP-Regret, a new metric for repeated games with adaptive opponents.
- RP-Regret is fundamentally a nonconvex problem, and the paper proposes three algorithms to address it.
- When all players minimize RP-Regret, they can learn a subgame-perfect equilibrium.
- On the Stag-Hunt game, the methods converge to more cooperative and higher-utility solutions than existing approaches.
- It identifies necessary conditions on the memory of comparison strategies and opponent strategies.
Paper links
External research summaries. These are not HDATF publications or measured product results.