Regret Minimization with Adaptive Opponents in Repeated Games

Published
Source
arXiv
Paper number
329
Field
Machine Learning
arXiv ID
2606.06486

Key points

  • It introduces RP-Regret, a new metric for repeated games with adaptive opponents.
  • RP-Regret is fundamentally a nonconvex problem, and the paper proposes three algorithms to address it.
  • When all players minimize RP-Regret, they can learn a subgame-perfect equilibrium.
  • On the Stag-Hunt game, the methods converge to more cooperative and higher-utility solutions than existing approaches.
  • It identifies necessary conditions on the memory of comparison strategies and opponent strategies.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)