Fetching the paper…
Reading the bibliography…
There exist many algorithms for learning how to play repeated bimatrix games.
In: Activity Analysis of Production and Allocation
Brown, G.: 1951, ‘Iterative solution of games by ficticious play’ · 1951
Earlier work this paper cites.
In: Journal of the Society for Industrial and Applied Mathematics
Lemke, C. and J. Howson: 1964, ‘Equilibrium points of bimatrix games.’ · 1964
Earlier work this paper cites.
Univeristy of Michigan Press
Rapoport, A., M. Guyer, and D. Gordon: 1976, The 2x2 Game · 1976
Earlier work this paper cites.
Operations Research
Heidelberger, P. and P. Lewis: 1984, ‘Quantile estimation in dependent sequences’ · 1984
Earlier work this paper cites.
In: L. Davis (ed.): Genetic Algorithms and Simulated Annealing
Axelrod, R.: 1987, ‘The Evolution of Strategies in the Iterated Prisoner’s Dilemma’ · 1987
Earlier work this paper cites.
Machine Learning
Watkins, C. and P. Dayan: 1992, ‘Q-learning: technical note’ · 1992
Earlier work this paper cites.
Games and Economic Behavior
Fudenberg, D. and D. M. Kreps: 1993, ‘Learning Mixed Equilibria’ · 1993
Earlier work this paper cites.
In: ICML 11
Littman, M.: 1994, ‘Markov games as a framework for multi-agent reinforcement learning’ · 1994
Earlier work this paper cites.
MIT Press
Osborne, M. and A. Rubinstein: 1994, A Course in Game Theory · 1994
Earlier work this paper cites.
Games and Economic Behavior
Monderer, D. and A. Sela: 1996, ‘A 2 × \times 2 game without the fictitious play property’ · 1996
Earlier work this paper cites.
Journal of Economic Theory
Monderer, D. and L. Shapley: 1996, ‘Fictitious play property for games with identical interests’ · 1996
Earlier work this paper cites.
In: AAAI 4
Claus, C. and C. Boutilier: 1997, ‘The dynamics of reinforcement learning in cooperative multiagent systems’ · 1997
Earlier work this paper cites.
In: ICML 15
Hu, J. and M. Wellman: 1998, ‘Multiagent reinforcement learning: theoretical framework and an algorithm’ · 1998
Earlier work this paper cites.
Cambridge, Massachusetts: The MIT Press
Sutton, R. and A. Barto: 1999, Reinforcement Learning, An Introduction · 1999
Cited alongside, same era.
In: UAI 16
Singh, S., M. Kearns, and Y. Mansour: 2000, ‘Nash convergence of gradient dynamics in general-sum games’ · 2000
Cited alongside, same era.
In: IJCAI 17
Bowling, M. and M. Veloso: 2001, ‘Rational and convergent learning in stochastic games’ · 2001
Cited alongside, same era.
In: ICML 18
Littman, M.: 2001, ‘Friend-or-foe Q-learning in general-sum games’ · 2001
Cited alongside, same era.
Artificial Intelligence
Bowling, M. H. and M. M. Veloso: 2002, ‘Multiagent learning using a variable learning rate’ · 2002
Cited alongside, same era.
In: ICML 20
Conitzer, V. and T. Sandholm: 2003, ‘AWESOME: A General Multiagent Learning Algorithm that Converges in Self-Play and Learns a Best Response Against Stationary Opponents’ · 2003
Cited alongside, same era.
In: AAMAS 3
Nudelman, E., J. Wortman, K. Leyton-Brown, and Y. Shoham: 2004, ‘Run the GAMUT: a comprehensive approach to evaluating game-theoretic algorithms’ · 2004
Later among the works it cites.
In: NIPS 16
Tesauro, G.: 2004, ‘Extending Q-learning to general adaptive multi-agent systems’ · 2004
Later among the works it cites.
Master’s thesis, University of British Columbia, Vancouver, Canada
Lipson, A.: 2005, ‘An empirical evaluation of multiagent learning algorithms’ · 2005
Later among the works it cites.
In: NIPS
Powers, R. and Y. Shoham: 2005, ‘New criteria and a new algorithm for learning in multi-agent systems’ · 2005
Later among the works it cites.
In: AAMAS
Vu, T., R. Powers, and Y. Shoham: 2005, ‘Learning against multiple opponents’ · 2005
Later among the works it cites.
In: AAMAS ’06
Banerjee, B. and J. Peng: 2006, ‘RV: a unifying approach to performance and convergence in online multiagent learning’ · 2006
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Govindan, S. and R. Wilson: 2003, ‘A global newton method to compute Nash equilibria’ · 2003
Cited alongside, same era.
In: ICML 20
Greenwald, A. and K. Hall: 2003, ‘Correlated-Q learning’ · 2003
Cited alongside, same era.
Journal of Machine Learning Research
Hu, J. and M. P. Wellman: 2003, ‘Nash Q-learning for general-sum stochastic games’ · 2003
Cited alongside, same era.
Hoboken, New Jersey: John Wiley & Sons
Spall, J. C.: 2003, Introduction to Stochastic Search and Optimization: Estimation, Simulation and Control · 2003
Cited alongside, same era.
In: ICML’03
Zinkevich, M.: 2003, ‘Online convex programming and generalized infinitesimal gradient ascent’ · 2003
Cited alongside, same era.
In: AAAI 11
Banerjee, B. and J. Peng: 2004, ‘Performance bounded reinforcement learning in strategic interactions’ · 2004
Cited alongside, same era.
R Foundation for Statistical Computing, Vienna, Austria
R Development Core Team: 2006, ‘R: a language and environment for statistical computing’ · 2006
Later among the works it cites.
Journal of Artificial Societies and Social Simulation
Airiau, S., S. Saha, and S. Sen: 2007, ‘Evolutionary Tournament-Based Comparison of Learning and Non-Learning Algorithms for Iterated Games’ · 2007
Later among the works it cites.
Machine Learning
Conitzer, V. and T. Sandholm: 2007, ‘AWESOME: A General Multiagent Learning Algorithm that Converges in Self-Play and Learns a Best Response Against Stationary Opponents’ · 2007
Later among the works it cites.
Artificial Intelligence
Sandholm, T.: 2007, ‘Perspectives on multiagent learning’ · 2007
Later among the works it cites.
Artificial Intelligence
Shoham, Y., R. Powers, and T. Grenager: 2007, ‘If multi-agent learning is the answer, what is the question?’ · 2007
Later among the works it cites.
New York: Cambridge University Press
Shoham, Y. and K. Leyton-Brown: 2008, Multiagent Systems: Algorithmic, Game-Theoretic, and Logical Foundations · 2008
Later among the works it cites.