Fetching the paper…
Reading the bibliography…
MOBA games, e.g., Honor of Kings, League of Legends, and Dota 2, pose grand challenges to AI systems such as multi-agent, enormous state-action space, complex action control, etc.
An analysis of alpha-beta pruning
D. E. Knuth and R. W. Moore · 1975
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Exact binomial confidence interval for proportions
J. T. Morisette and S. Khorram · 1998
Earlier work this paper cites.
Actor-critic algorithms
V. R. Konda and J. N. Tsitsiklis · 2000
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
R. Coulom · 2006
Earlier work this paper cites.
Bandit based monte-carlo planning
L. Kocsis and C. Szepesvári · 2006
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
L. Bu, R. Babu, B. De Schutter, et al · 2008
Earlier work this paper cites.
Whole-history rating: A bayesian rating system for players of time-varying strength
R. Coulom · 2008
Earlier work this paper cites.
Curriculum learning
Y. Bengio, J. Louradour, R. Collobert, and J. Weston · 2009
Earlier work this paper cites.
A survey of real-time strategy game ai research and competition in starcraft
S. Ontanón, G. Synnaeve, A. Uriarte, F. Richoux, D. Churchill, and M. Preuss · 2013
Earlier work this paper cites.
A review of real-time strategy game ai
G. Robertson and I. Watson · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
A. A. Rusu, S. G. Colmenarejo, C. Gulcehre, G. Desjardins, J. Kirkpatrick, R. Pascanu, V. Mnih, K. Kavukcuoglu, and R. Hadsell · 2015
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel · 2015
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Cited alongside, same era.
The game imitation: Deep supervised convolutional networks for quick video game ai
Z. Chen and D. Yi · 2017
Cited alongside, same era.
A survey of learning in multiagent environments: Dealing with non-stationarity
Feedback-based tree search for reinforcement learning
D. Jiang, E. Ekwedike, and H. Liu · 2018
Later among the works it cites.
Exponentially weighted imitation learning for batched historical data
Q. Wang, J. Xiong, L. Han, H. Liu, and T. Zhang · 2018
Later among the works it cites.
Dota 2 with large scale deep reinforcement learning
C. Berner, G. Brockman, B. Chan, V. Cheung, P. Dębiak, C. Dennison, D. Farhi, Q. Fischer, S. Hashme, C. Hesse, et al · 2019
Later among the works it cites.
Distilling policy distillation
W. M. Czarnecki, R. Pascanu, S. Osindero, S. Jayakumar, G. Swirszcz, and M. Jaderberg · 2019
Later among the works it cites.
Marginal policy gradients: A unified family of estimators for bounded action spaces with applications
C. Eisenach, H. Yang, J. Liu, and H. Liu · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Hernandez-Leal, M. Kaisers, T. Baarslag, and E. M. de Cote · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
V. d. N. Silva and L. Chaimowicz · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Cited alongside, same era.
Hybrid reward architecture for reinforcement learning
H. Van Seijen, M. Fatemi, J. Romoff, R. Laroche, T. Barnes, and J. Tsang · 2017
Cited alongside, same era.
Starcraft ii: A new challenge for reinforcement learning
O. Vinyals, T. Ewalds, S. Bartunov, P. Georgiev, A. S. Vezhnevets, M. Yeo, A. Makhzani, H. Küttler, J. Agapiou, J. Schrittwieser, et al · 2017
Cited alongside, same era.
Superhuman ai for heads-up no-limit poker: Libratus beats top professionals
N. Brown and T. Sandholm · 2018
Cited alongside, same era.
The art of drafting: a team-oriented hero recommendation system for multiplayer online battle arena games
Z. Chen, T.-H. D. Nguyen, Y. Xu, C. Amato, S. Cooper, Y. Sun, and M. S. El-Nasr · 2018
Cited alongside, same era.
L. Espeholt, R. Marinier, P. Stanczyk, K. Wang, and M. Michalski · 2019
Later among the works it cites.
Human-level performance in 3d multiplayer games with population-based reinforcement learning
M. Jaderberg, W. M. Czarnecki, I. Dunning, L. Marris, G. Lever, A. G. Castañeda, C. Beattie, N. C. Rabinowitz, A. S. Morcos, A. Ruderman, et al · 2019
Later among the works it cites.
Openai five
OpenAI · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, et al · 2019
Later among the works it cites.
Hierarchical macro strategy model for moba game ai
B. Wu · 2019
Later among the works it cites.
Supervised learning achieves human-level performance in moba games: A case study of honor of kings
D. Ye, G. Chen, P. Zhao, F. Qiu, B. Yuan, W. Zhang, S. Chen, M. Sun, X. Li, S. Li, et al · 2020
Closest in time.
Mastering complex control in moba games with deep reinforcement learning
D. Ye, Z. Liu, M. Sun, B. Shi, P. Zhao, H. Wu, H. Yu, S. Yang, X. Wu, Q. Guo, et al · 2020
Closest in time.