Fetching the paper…
Reading the bibliography…
Recently, AlphaZero has achieved landmark results in deep reinforcement learning, by providing a single self-play architecture that learned three different games at super human level.
Gobang ist pspace-vollständig
Stefan Reisch · 1980
Earlier work this paper cites.
A knowledge-based approach of connect-four
L. Victor Allis · 1988
Earlier work this paper cites.
The othello game on an n*n board is pspace-complete
Shigeki Iwata and Takumi Kasai · 1994
Earlier work this paper cites.
Temporal difference learning and td-gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
The othello match of the year: Takeshi murakami vs. logistello
Michael Buro · 1997
Earlier work this paper cites.
New self-play results in computer chess
Ernst A Heinz · 2000
Earlier work this paper cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2004
Earlier work this paper cites.
Coevolution versus self-play temporal difference learning for acquiring position evaluation in small-board go
Thomas Philip Runarsson and Simon M Lucas · 2005
Earlier work this paper cites.
Observing the evolution of neural networks learning to play the game of othello
Siang Yew Chong, Mei K. Tan, and Jonathon David White · 2005
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
Rémi Coulom · 2006
Earlier work this paper cites.
Combining online and offline knowledge in uct
Sylvain Gelly and David Silver · 2007
Earlier work this paper cites.
General game learning using knowledge transfer
Bikramjit Banerjee and Peter Stone · 2007
Earlier work this paper cites.
Monte-carlo tree search: A new framework for game ai
Guillaume Chaslot, Sander Bakkes, Istvan Szita, and Pieter Spronck · 2008
Earlier work this paper cites.
Whole-history rating: A bayesian rating system for players of time-varying strength
Rémi Coulom · 2008
Earlier work this paper cites.
Self-play and using an expert to learn to play backgammon with temporal difference learning
Marco A Wiering et al · 2010
Earlier work this paper cites.
Monte-carlo tree search and rapid action value estimation in computer go
Sylvain Gelly and David Silver · 2011
Cited alongside, same era.
Multi-armed bandits with episode context
Christopher D Rosin · 2011
Cited alongside, same era.
A survey of monte carlo tree search methods
Cameron Browne, Edward Jack Powley, Daniel Whitehouse, Simon M. Lucas, Peter I. Cowling, Philipp Rohlfshagen, Stephen Tavener, Diego Perez Liebana, Spyridon Samothrakis, and Simon Colton · 2012
Cited alongside, same era.
Design of evaluation-function for computer gobang game system [j]
ML Zhang, Jun Wu, and FZ Li · 2012
Cited alongside, same era.
Reinforcement learning in the game of othello: learning against a fixed opponent and learning from self-play
Michiel Van Der Ree and Marco Wiering · 2013
Cited alongside, same era.
Rolling horizon evolution versus tree search for navigation in single-player real-time games
When doctors meet with alphago: potential application of machine learning to clinical medicine
Zhongheng Zhang · 2016
Later among the works it cites.
Rolling horizon coevolutionary planning for two-player video games
Jialin Liu, Diego Perez Liebana, and Simon M. Lucas · 2016
Later among the works it cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Later among the works it cites.
Playing multi-action adversarial games: Online evolutionary planning versus tree search
Niels Justesen, Tobias Mahlmann, Sebastian Risi, and Julian Togelius · 2017
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Diego Perez, Spyridon Samothrakis, Simon Lucas, and Philipp Rohlfshagen · 2013
Cited alongside, same era.
Combining simulated annealing and monte carlo tree search for expression simplification
Ben Ruijl, Jos Vermaseren, Aske Plaat, and Jaap van den Herik · 2014
Cited alongside, same era.
Temporal difference learning with eligibility traces for the game connect four
Markus Thill, Samineh Bagheri, Patrick Koch, and Wolfgang Konen · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin A. Riedmiller, Andreas Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
Training deep convolutional neural networks to play go
Christopher Clark and Amos Storkey · 2015
Cited alongside, same era.
Alphazero general
Surag Nair · 2018
Later among the works it cites.
Monte carlo q-learning for general game playing
Hui Wang, Michael Emmerich, and Aske Plaat · 2018
Later among the works it cites.
Assessing the potential of classical q-learning in general game playing
Hui Wang, Michael Emmerich, and Aske Plaat · 2018
Later among the works it cites.
The 2016 two-player GVGAI competition
Raluca D. Gaina, Adrien Couëtoux, Dennis J. N. J. Soemers, Mark H. M. Winands, Tom Vodopivec, Florian Kirchgeßner, Jialin Liu, Simon M. Lucas, and Diego Pérez-Liébana · 2018
Later among the works it cites.
Accelerating self-play learning in go
David J Wu · 2019
Later among the works it cites.
Alternative loss functions in alphazero-like self-play
Hui Wang, Michael Emmerich, Mike Preuss, and Aske Plaat · 2019
Later among the works it cites.
Hyper-parameter sweep on alphazero general
Hui Wang, Michael Emmerich, Mike Preuss, and Aske Plaat · 2019
Later among the works it cites.
Learning to play—reinforcement learning and games, 2020
Aske Plaat · 2020
Closest in time.
Rolling horizon evolutionary algorithms for general video game playing, 2020
Raluca D. Gaina, Sam Devlin, Simon M. Lucas, and Diego Perez-Liebana · 2020
Closest in time.
Analysis of hyper-parameters for small games: Iterations or epochs in self-play?
Hui Wang, Michael Emmerich, Mike Preuss, and Aske Plaat · 2020
Closest in time.