Fetching the paper…
Reading the bibliography…
Existing multi-agent reinforcement learning methods are limited typically to a small number of agents.
Beitrag zur theorie des ferromagnetismus
Ising, E · 1925
Earlier work this paper cites.
Stochastic games
Shapley, L. S · 1953
Earlier work this paper cites.
Equilibrium in a stochastic n n -person game
Fink, A. M. et al · 1964
Earlier work this paper cites.
Equilibrium points of bimatrix games
Lemke, C. E. and Howson, Jr, J. T · 1964
Earlier work this paper cites.
Phase transitions and critical phenomena
Stanley, H. E · 1971
Earlier work this paper cites.
Introductory functional analysis with applications , volume 1
Kreyszig, E · 1978
Earlier work this paper cites.
Stochastic Dynamic Programming: successive approximations and nearly optimal strategies for Markov decision processes and Markov games
van der Wal, J., van der Wal, J., van der Wal, J., Mathématicien, P.-B., van der Wal, J., and Mathematician, N · 1981
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P · 1992
Earlier work this paper cites.
Monte carlo simulation in statistical physics
Binder, K., Heermann, D., Roelofs, L., Mallinckrodt, A. J., and McKay, S · 1993
Earlier work this paper cites.
The statistical mechanics of strategic interaction
Blume, L. E · 1993
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Tan, M · 1993
Earlier work this paper cites.
Convergence of stochastic iterative dynamic programming algorithms
Jaakkola, T., Jordan, M. I., and Singh, S. P · 1994
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Littman, M. L · 1994
Earlier work this paper cites.
Envisioning stock trading where the brokers are bots
Troy, C. A · 1997
Earlier work this paper cites.
A unified analysis of value-function-based reinforcement-learning algorithms
Szepesvári, C. and Littman, M. L · 1999
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N · 2000
Earlier work this paper cites.
Convergence results for single-step on-policy reinforcement-learning algorithms
Singh, S., Jaakkola, T., Littman, M. L., and Szepesvári, C · 2000
Earlier work this paper cites.
Rational and convergent learning in stochastic games
Bowling, M. and Veloso, M · 2001
Cited alongside, same era.
Friend-or-foe q-learning in general-sum games
Littman, M. L · 2001
Cited alongside, same era.
Multiagent learning using a variable learning rate
Bowling, M. and Veloso, M · 2002
Cited alongside, same era.
Reinforcement learning of coordination in cooperative multi-agent systems
Kapetanakis, S. and Kudenko, D · 2002
Cited alongside, same era.
Nash q-learning for general-sum stochastic games
Hu, J. and Wellman, M. P · 2003
Cited alongside, same era.
A polynomial-time nash equilibrium algorithm for repeated games
Littman, M. L. and Stone, P · 2005
Cited alongside, same era.
Weighted sup-norm contractions in dynamic programming: A review and some new applications
Bertsekas, D. P · 2012
Later among the works it cites.
Independent reinforcement learners in cooperative markov games: a survey regarding coordination problems
Matignon, L., Laurent, G. J., and Le Fort-Piat, N · 2012
Later among the works it cites.
Factorization machines with libfm
Rendle, S · 2012
Later among the works it cites.
Exploiting structure and agent-centric rewards to promote coordination in large multiagent systems
HolmesParker, C., Taylor, M., Zhan, Y., and Tumer, K · 2014
Later among the works it cites.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M · 2014
Later among the works it cites.
Counterfactual exploration for improving multiagent learning
Colby, M. K., Kharaghani, S., HolmesParker, C., and Tumer, K · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cooperative multi-agent learning: The state of the art
Panait, L. and Luke, S · 2005
Cited alongside, same era.
Large population stochastic dynamic games: closed-loop mckean-vlasov systems and the nash certainty equivalence principle
Huang, M., Malhamé, R. P., Caines, P. E., et al · 2006
Cited alongside, same era.
Oblivious equilibrium: A mean field approximation for large-scale dynamic games
Weintraub, G. Y., Benkard, L., and Van Roy, B · 2006
Cited alongside, same era.
Learning to rank: from pairwise approach to listwise approach
Cao, Z., Qin, T., Liu, T.-Y., Tsai, M.-F., and Li, H · 2007
Cited alongside, same era.
Mean field games
Lasry, J.-M. and Lions, P.-L · 2007
Cited alongside, same era.
A comprehensive survey of multiagent reinforcement learning
Busoniu, L., Babuska, R., and De Schutter, B · 2008
Cited alongside, same era.
Later among the works it cites.
Analysis of game bot’s behavioral characteristics in social interaction networks of mmorpg
Jeong, S. H., Kang, A. R., and Kim, H. K · 2015
Later among the works it cites.
Opponent modeling in deep reinforcement learning
He, H. and Boyd-Graber, J. L · 2016
Later among the works it cites.
Learning with opponent-learning awareness
Foerster, J. N., Chen, R. Y., Al-Shedivat, M., Whiteson, S., Abbeel, P., and Mordatch, I · 2017
Later among the works it cites.
Cooperative multi-agent control using deep reinforcement learning
Gupta, J. K., Egorov, M., and Kochenderfer, M · 2017
Later among the works it cites.
Multiagent bidirectionally-coordinated nets for learning to play starcraft combat games
Peng, P., Yuan, Q., Wen, Y., Yang, Y., Tang, Z., Long, H., and Wang, J · 2017
Later among the works it cites.
Display advertising with real-time bidding (rtb) and behavioural targeting
Wang, J., Zhang, W., Yuan, S., et al · 2017
Later among the works it cites.
Deep mean field games for learning optimal behavior policy of large populations
Yang, J., Ye, X., Trivedi, R., Xu, H., and Zha, H · 2017
Later among the works it cites.
Counterfactual multi-agent policy gradients
Foerster, J. N., Farquhar, G., Afouras, T., Nardelli, N., and Whiteson, S · 2018
Closest in time.
Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, New Orleans, Louisiana, USA, February 2-7, 2018 , 2018. AAAI Press
McIlraith, S. A. and Weinberger, K. Q. (eds.) · 2018
Closest in time.
Magent: A many-agent reinforcement learning platform for artificial collective intelligence
Zheng, L., Yang, J., Cai, H., Zhou, M., Zhang, W., Wang, J., and Yu, Y · 2018
Closest in time.