Fetching the paper…
Reading the bibliography…
In reinforcement learning (RL), the term self-play describes a kind of multi-agent learning (MAL) that deploys an algorithm against copies of itself to test compatibility in various stochastic environments.
The bargaining problem
J. Nash · 1950
Earlier work this paper cites.
1. Some Topics in Two-Person Games
L. S. Shapley · 1964
Earlier work this paper cites.
Proportional solutions to bargaining situations: Intertemporal utility comparisons
E. Kalai · 1977
Earlier work this paper cites.
Rational learning leads to nash equilibrium
E. Kalai and E. Lehrer · 1993
Earlier work this paper cites.
A Course in Game Theory
M. J. Osborne and A. Rubinstein · 1994
Earlier work this paper cites.
R-max - a general polynomial time algorithm for near-optimal reinforcement learning
R. I. Brafman and M. Tennenholtz · 2002
Earlier work this paper cites.
Learning ϵ \epsilon -pareto efficient solutions with minimal knowledge requirements using satisficing
J. W. Crandall and M. A. Goodrich · 2004
Earlier work this paper cites.
New criteria and a new algorithm for learning in multi-agent systems
R. Powers and Y. Shoham · 2004
Cited alongside, same era.
From external to internal regret
A. Blum and Y. Mansour · 2005
Cited alongside, same era.
Learning against opponents with bounded memory
R. Powers and Y. Shoham · 2005
Cited alongside, same era.
Awesome: A general multiagent learning algorithm that converges in self-play and learns a best response against stationary opponents
V. Conitzer and T. Sandholm · 2006
Cited alongside, same era.
If multi-agent learning is the answer, what is the question?
Y. Shoham, R. Powers, and T. Grenager · 2007
Cited alongside, same era.
A polynomial-time nash equilibrium algorithm for repeated stochastic games
E. M. de Cote and M. L. Littman · 2008
Cited alongside, same era.
Learning to compete, coordinate, and cooperate in repeated games using reinforcement learning
J. W. Crandall and M. A. Goodrich · 2010
Later among the works it cites.
Online bandit learning against an adaptive adversary: from regret to policy regret
R. Arora, O. Dekel, and A. Tewari · 2012
Later among the works it cites.
Towards minimizing disappointment in repeated games
J. W. Crandall · 2014
Later among the works it cites.
Foolproof cooperative learning
A. Jacq, J. Perolat, M. Geist, and O. Pietquin · 2020
Later among the works it cites.
A novel individually rational objective in multi-agent multi-armed bandits: Algorithms and regret bounds
A. C. Tossou, C. Dimitrakakis, J. Rzepecki, and K. Hofmann · 2020
Later among the works it cites.
A sharp analysis of model based reinforcement learning with self-play
Q. Liu, T. Yu, Y. Bai, and C. Jin · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Convergence, targeted optimality and safety in multiagent learning
D. Chakraborty and P. Stone · 2010
Cited alongside, same era.
Closest in time.