Fetching the paper…
Reading the bibliography…
Monte Carlo Tree Search (MCTS) algorithms perform simulation-based search to improve policies online.
Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning
R. J. Williams · 1992
Earlier work this paper cites.
On-line policy improvement using monte-carlo search
Gerald Tesauro and Gregory R Galperin · 1997
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2002
Earlier work this paper cites.
Reinforcement learning as classification: Leveraging modern classifiers
Michail G Lagoudakis and Ronald Parr · 2003
Earlier work this paper cites.
Bayeselo
R Coulom · 2005
Earlier work this paper cites.
Bandit Based Monte-Carlo Planning
L. Kocsis and C. Szepesvári · 2006
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
Rémi Coulom · 2006
Earlier work this paper cites.
Reinforcement learning and simulation-based search
David Silver · 2009
Earlier work this paper cites.
On the scalability of parallel uct
Richard B Segal · 2010
Earlier work this paper cites.
Monte Carlo Tree Search in Hex
B. Arneson, R. Hayward, and P. Hednerson · 2010
Earlier work this paper cites.
Continuous upper confidence trees
Adrien Couëtoux, Jean-Baptiste Hoock, Nataliya Sokolovska, Olivier Teytaud, and Nicolas Bonnard · 2011
Earlier work this paper cites.
Multi-Armed Bandits with Episode Context
C. D. Rosin · 2011
Cited alongside, same era.
A Survey of Monte Carlo Tree Search Methods
C. B. Browne, E. Powley, D. Whitehouse, S. M. Lucas, P. I. Cowling, P. Rohlfshagen, S. Tavener, D. Perez, S. Samothrakis, and S. Colton · 2012
Cited alongside, same era.
Mohex 2.0: a pattern-based mcts hex player
Shih-Chieh Huang, Broderick Arneson, Ryan B Hayward, Martin Müller, and Jakub Pawlewicz · 2013
Cited alongside, same era.
Scalable parallel dfpn search
Jakub Pawlewicz and Ryan B Hayward · 2013
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
Approximate modified policy iteration and its application to the game of tetris
Bruno Scherrer, Mohammad Ghavamzadeh, Victor Gabillon, Boris Lesner, and Matthieu Geist · 2015
Thinking fast and slow with deep learning and tree search
Thomas Anthony, Zheng Tian, and David Barber · 2017
Later among the works it cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2017
Later among the works it cites.
Ray: A distributed framework for emerging ai applications
Philipp Moritz, Robert Nishihara, Stephanie Wang, Alexey Tumanov, Richard Liaw, Eric Liang, William Paul, Michael I Jordan, and Ion Stoica · 2017
Later among the works it cites.
Reinforcement learning for creating evaluation function using convolutional neural network in hex
Kei Takada, Hiroyuki Iizuka, and Masahito Yamamoto · 2017
Later among the works it cites.
Move prediction using deep convolutional neural networks in hex
Chao Gao, Ryan B Hayward, and Martin Müller · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Adaptive playouts in monte-carlo tree search with policy-gradient reinforcement learning
Tobias Graf and Marco Platzner · 2015
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
NeuroHex: A Deep Q-learning Hex Agent
K. Young, R. Hayward, and G. Vasan · 2016
Cited alongside, same era.
Mastering the Game of Go without Human Knowledge
D. Silver et al · 2017
Cited alongside, same era.
Later among the works it cites.
Distral: Robust multitask reinforcement learning
Yee Teh, Victor Bapst, Wojciech M Czarnecki, John Quan, James Kirkpatrick, Raia Hadsell, Nicolas Heess, and Razvan Pascanu · 2017
Later among the works it cites.
Divide-and-conquer reinforcement learning
Dibya Ghosh, Avi Singh, Aravind Rajeswaran, Vikash Kumar, and Sergey Levine · 2017
Later among the works it cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Adversarial policy gradient for alternating markov games
Chao Gao, Martin Mueller, and Ryan Hayward · 2018
Later among the works it cites.