Learning with Bandit Feedback in Potential Games

Amélie HeliouJohanne CohenPanayotis Mertikopoulos

This paper examines the equilibrium convergence properties of no-regret learning with exponential weights in potential games. To establish convergence with minimal information requirements on the players' side, we focus on two frameworks: the semi-bandit case (where players have access to a noisy estimate of their payoff vectors, including strategies they did not play), and the bandit case (where players are only able to observe their in-game, realized payoffs)... (read more)

PDF Abstract