Fetching the paper…
Reading the bibliography…
A core novelty of Alpha Zero is the interleaving of tree search and deep learning, which has proven very successful in board games like Chess, Shogi and Go.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Efficient selectivity and backup operators in Monte-Carlo tree search
Coulom, R. (2006) · 2006
Earlier work this paper cites.
Bandit based monte-carlo planning
Kocsis, L. and Szepesvári, C. (2006) · 2006
Earlier work this paper cites.
Computing elo ratings of move patterns in the game of go
Coulom, R. (2007) · 2007
Earlier work this paper cites.
Progressive Strategies For Monte-Carlo Tree Search
Chaslot, G. M. J., Winands, M. H., Van Den Herik, H. J., Uiterwijk, J. W., Bouzy, B., et al. (2008) · 2008
Earlier work this paper cites.
Continuous upper confidence trees
Couëtoux, A., Hoock, J.-B., Sokolovska, N., Teytaud, O., and Bonnard, N. (2011) · 2011
Earlier work this paper cites.
Multi-armed bandits with episode context
Rosin, C. D. (2011) · 2011
Cited alongside, same era.
A survey of monte carlo tree search methods
Browne, C. B., Powley, E., Whitehouse, D., Lucas, S. M., Cowling, P. I., Rohlfshagen, P., Tavener, S., Perez, D., Samothrakis, S., and Colton, S. (2012) · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y. (2012) · 2012
Cited alongside, same era.
Handbook of differential entropy
Michalowicz, J. V., Nichols, J. M., and Bucholtz, F. (2013) · 2013
Cited alongside, same era.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015) · 2015
Cited alongside, same era.
TensorFlow: A System for Large-Scale Machine Learning
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., et al. (2016) · 2016
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. (2016) · 2016
Later among the works it cites.
Efficient exploration with Double Uncertain Value Networks
Moerland, T. M., Broekens, J., and Jonker, C. M. (2017) · 2017
Later among the works it cites.
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. (2018) · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al. (2017a)
Cited in the paper.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al. (2017b)
Cited in the paper.
Moerland, T. M., Broekens, J., Plaat, A., and Jonker, C. M. (2018) · 2018
Closest in time.
Reinforcement learning: An Introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Closest in time.