Fetching the paper…
Reading the bibliography…
The landmark achievements of AlphaGo Zero have created great research interest into self-play in reinforcement learning.
Reisch, S. Gobang ist PSPACE-vollständig. Acta Informatica 13, 59¨C66 (1980)
1980
Earlier work this paper cites.
Allis V. A knowledge-based approach of Connect-Four-the game is solved: White wins. 1988
1988
Earlier work this paper cites.
Iwata S, Kasai T. The Othello game on an n × \times n board is PSPACE-complete. Theoretical Computer Science. 123
1994
Earlier work this paper cites.
Caruana R. Multitask learning. Machine learning, 1997, 28(1): 41-75
1997
Earlier work this paper cites.
Buro M. The Othello match of the year: Takeshi Murakami vs. Logistello. ICGA Journal, 1997, 20(3): 189-193
1997
Earlier work this paper cites.
Heinz E A: New self-play results in computer chess. International Conference on Computers and Games. Springer, Berlin, Heidelberg. pp. 262–276 (2000)
2000
Earlier work this paper cites.
Birattari M, Stützle T, Paquete L, et al. A racing algorithm for configuring metaheuristics. Proceedings of the 4th Annual Conference on Genetic and Evolutionary Computation. Morgan Kaufmann Publishers Inc. 11-18 (2002)
2002
Earlier work this paper cites.
Chong S Y, Tan M K, White J D. Observing the evolution of neural networks learning to play the game of Othello. IEEE Transactions on Evolutionary Computation, 2005, 9(3): 240-251
2005
Earlier work this paper cites.
Banerjee B, Stone P. General Game Learning Using Knowledge Transfer. IJCAI. 2007: 672-677
2007
Earlier work this paper cites.
Coulom R. Whole-history rating: A Bayesian rating system for players of time-varying strength. International Conference on Computers and Games. Springer, Berlin, Heidelberg, 113–124, 2008
2008
Earlier work this paper cites.
Wiering M A: Self-Play and Using an Expert to Learn to Play Backgammon with Temporal Difference Learning. Journal of Intelligent Learning Systems and Applications 2
2010
Earlier work this paper cites.
Hutter F, Hoos H H, Leyton-Brown K: Sequential model-based optimization for general algorithm configuration. International Conference on Learning and Intelligent Optimization. Springer, Berlin, Heidelberg, pp. 507–523 (2011)
2011
Earlier work this paper cites.
Browne C B, Powley E, Whitehouse D, et al: A survey of monte carlo tree search methods. IEEE Transactions on Computational Intelligence and AI in games 4
2012
Cited alongside, same era.
Zhang M L, Wu J, Li F Z. Design of evaluation-function for computer Gobang game system [J][J]. Journal of Computer Applications, 2012, 7: 051
2012
Cited alongside, same era.
Van Der Ree M, Wiering M: Reinforcement learning in the game of Othello: Learning against a fixed opponent and learning from self-play. In Adaptive Dynamic Programming And Reinforcement Learning. pp. 108–115 (2013)
2013
Cited alongside, same era.
B Ruijl, J Vermaseren, A Plaat, J Herik: Combining Simulated Annealing and Monte Carlo Tree Search for Expression Simplification. In: Béatrice Duval, H. Jaap van den Herik, Stéphane Loiseau, Joaquim Filipe. Proceedings of the 6th International Conference on Agents and Artificial Intelligence 2014, vol. 1, pp. 724–731. SciTePress, Setúbal, Portugal (2014)
2014
Cited alongside, same era.
Zhang Z: When doctors meet with AlphaGo: potential application of machine learning to clinical medicine. Annals of translational medicine 4
2016
Later among the works it cites.
Silver D, Schrittwieser J, Simonyan K, et al: Mastering the game of go without human knowledge. Nature 550
2017
Later among the works it cites.
Silver D, Hubert T, Schrittwieser J, et al. A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play. Science, 2018, 362(6419): 1140-1144
2018
Later among the works it cites.
N. Surag, https://github.com/suragnair/alpha-zero-general, 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Thill M, Bagheri S, Koch P, et al. Temporal difference learning with eligibility traces for the game connect four. 2014 IEEE Conference on Computational Intelligence and Games. IEEE, 2014: 1-8
2014
Cited alongside, same era.
Kingma D P, Ba J: Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014)
2014
Cited alongside, same era.
Srivastava N, Hinton G, Krizhevsky A, et al: Dropout: a simple way to prevent neural networks from overfitting. The Journal of Machine Learning Research. 15
2014
Cited alongside, same era.
Schmidhuber J: Deep learning in neural networks: An overview. Neural networks 61
2015
Cited alongside, same era.
Clark C, Storkey A. Training deep convolutional neural networks to play go. International Conference on Machine Learning. pp. 1766–1774 (2015)
2015
Cited alongside, same era.
Ioffe S, Szegedy C: Batch normalization: accelerating deep network training by reducing internal covariate shift. Proceedings of the 32nd International Conference on International Conference on Machine Learning-Volume 37. pp. 448–456 (2015)
2015
Cited alongside, same era.
Silver D, Huang A, Maddison C J, et al: Mastering the game of Go with deep neural networks and tree search. Nature 529
2016
Cited alongside, same era.
Tao J, Wu L, Hu X: Principle Analysis on AlphaGo and Perspective in Military Application of Artificial Intelligence. Journal of Command and Control 2
2016
Cited alongside, same era.
Mandai Y, Kaneko T. Alternative Multitask Training for Evaluation Functions in Game of Go. 2018 Conference on Technologies and Applications of Artificial Intelligence (TAAI). IEEE, 2018: 132-135
2018
Later among the works it cites.
Matsuzaki K, Kitamura N. Do evaluation functions really improve Monte-Carlo tree search?[J]. ICGA Journal, 2018 (Preprint): 1-11
2018
Later among the works it cites.
Matsuzaki K. Empirical Analysis of PUCT Algorithm with Evaluation Functions of Different Quality. 2018 Conference on Technologies and Applications of Artificial Intelligence (TAAI). IEEE, 2018: 142-147
2018
Later among the works it cites.
Wang H., Emmerich M., Plaat A. (2019) Assessing the Potential of Classical Q-learning in General Game Playing. In: Atzmueller M., Duivesteijn W. (eds) Artificial Intelligence. BNAIC 2018. Communications in Computer and Information Science, vol 1021. Springer, Cham
2018
Later among the works it cites.
Emmerich M T M, Deutz A H. A tutorial on multiobjective optimization: fundamentals and evolutionary methods. Natural computing, 2018, 17(3): 585-609
2018
Later among the works it cites.
Wang H, Emmerich M, Preuss M and Plaat A. Alternative Loss Functions in AlphaZero-like Self-play. 2019 IEEE Symposium Series on Computational Intelligence (SSCI), Xiamen, China, 155–162 (2019)
2019
Later among the works it cites.
Aske Plaat, Learning to Play: Reinforcement Learning and Games, Leiden, 2020, forthcoming
2020
Closest in time.