Fetching the paper…
Reading the bibliography…
Inspired by recent successes of Monte-Carlo tree search (MCTS) in a number of artificial intelligence (AI) application domains, we propose a model-based reinforcement learning (RL) technique that iteratively applies MCTS on batches of small, finite-horizon versions of the original infinite-horizon Markov decision process.
Convergence of discretization procedures in dynamic programming
Bertsekas, Dimitri P · 1975
Earlier work this paper cites.
Empirical processes: Theory and applications
Pollard, David · 1990
Earlier work this paper cites.
Decision theoretic generalizations of the PAC model for neural net and other learning applications
Haussler, David · 1992
Earlier work this paper cites.
Neuro-dynamic Programming
Bertsekas, Dimitri P and Tsitsiklis, John N · 1996
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, Andrew Y, Harada, Daishi, and Russell, Stuart · 1999
Earlier work this paper cites.
On the existence of fixed points for approximate value iteration and temporal-difference learning
De Farias, D Pucci and Van Roy, Benjamin · 2000
Earlier work this paper cites.
Deep blue
Campbell, Murray, Hoane Jr, A Joseph, and Hsu, Feng-hsiung · 2002
Earlier work this paper cites.
Nash Q-learning for general-sum stochastic games
Hu, Junling and Wellman, Michael P · 2003
Earlier work this paper cites.
Evaluation in go by a neural network using soft segmentation
Enzenberger, Markus · 2004
Earlier work this paper cites.
Relating reinforcement learning performance to classification performance
Langford, John and Zadrozny, Bianca · 2005
Earlier work this paper cites.
Monte-Carlo strategies for computer Go
Chaslot, Guillaume, Saito, Jahn-Takeshi, Uiterwijk, Jos WHM, Bouzy, Bruno, and van den Herik, H Jaap · 2006
Earlier work this paper cites.
Efficient selectivity and backup operators in Monte-Carlo tree search
Coulom, Rémi · 2006
Earlier work this paper cites.
Bandit based Monte-Carlo planning
Kocsis, Levente and Szepesvári, Csaba · 2006
Earlier work this paper cites.
Performance loss bounds for approximate value iteration with state aggregation
Van Roy, Benjamin · 2006
Earlier work this paper cites.
Combining online and offline knowledge in UCT
Gelly, Sylvain and Silver, David · 2007
Earlier work this paper cites.
Experiments with Monte Carlo Othello
Hingston, Philip and Masek, Martin · 2007
Cited alongside, same era.
Focus of attention in reinforcement learning
Li, Lihong, Bulitko, Vadim, and Greiner, Russell · 2007
Cited alongside, same era.
Performance bounds in l_p-norm for approximate value iteration
Munos, Rémi · 2007
Cited alongside, same era.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
Antos, András, Szepesvári, Csaba, and Munos, Rémi · 2008
Cited alongside, same era.
Monte-carlo tree search: A new framework for game AI
Chaslot, Guillaume, Bakkes, Sander, Szita, Istvan, and Spronck, Pieter · 2008
Cited alongside, same era.
Adaptative play in Texas hold’em poker
Maîtrepierre, Raphaël, Mary, Jérémie, and Munos, Rémi · 2008
Cited alongside, same era.
Deep learning for real-time Atari game play using offline Monte-Carlo tree search planning
Guo, Xiaoxiao, Singh, Satinder, Lee, Honglak, Lewis, Richard L, and Wang, Xiaoshi · 2014
Later among the works it cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, Martin L · 2014
Later among the works it cites.
Efficient sampling method for Monte Carlo tree search problem
Teraoka, Kazuki, Hatano, Kohei, and Takimoto, Eiji · 2014
Later among the works it cites.
Approximate modified policy iteration and its application to the game of tetris
Scherrer, Bruno, Ghavamzadeh, Mohammad, Gabillon, Victor, Lesner, Boris, and Geist, Matthieu · 2015
Later among the works it cites.
Al-Kanj, Lina, Powell, Warren B, and Bouzaiene-Ayari, Belgacem · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Finite-time bounds for fitted value iteration
Munos, Rémi and Szepesvári, Csaba · 2008
Cited alongside, same era.
Nested Monte-Carlo search
Cazenave, Tristan · 2009
Cited alongside, same era.
Combining UCT and nested Monte Carlo search for single-player general game playing
Méhat, Jean and Cazenave, Tristan · 2010
Cited alongside, same era.
Monte-carlo tree search and rapid action value estimation in computer Go
Gelly, Sylvain and Silver, David · 2011
Cited alongside, same era.
A survey of monte carlo tree search methods
Browne, Cameron B, Powley, Edward, Whitehouse, Daniel, Lucas, Simon M, Cowling, Peter I, Rohlfshagen, Philipp, Tavener, Stephen, Perez, Diego, Samothrakis, Spyridon, and Colton, Simon · 2012
Cited alongside, same era.
Approximation of markov decision processes with general state space
Dufour, François and Prieto-Rumeau, Tomás · 2012
Cited alongside, same era.
Empirical dynamic programming
Haskell, William B, Jain, Rahul, and Kalathil, Dileep · 2016
Later among the works it cites.
Analysis of classification-based policy iteration algorithms
Lazaric, Alessandro, Ghavamzadeh, Mohammad, and Munos, Rémi · 2016
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D · 2016
Later among the works it cites.
Thinking fast and slow with deep learning and tree search
Anthony, Thomas, Tian, Zheng, and Barber, David · 2017
Later among the works it cites.
An empirical dynamic programming algorithm for continuous MDPs
Haskell, William B, Jain, Rahul, Sharma, Hiteshi, and Yu, Pengqian · 2017
Later among the works it cites.
Monte carlo tree search with sampled information relaxation dual bounds
Jiang, Daniel R, Al-Kanj, Lina, and Powell, Warren B · 2017
Later among the works it cites.
Monte-Carlo tree search by best arm identification
Kaufmann, Emilie and Koolen, Wouter · 2017
Later among the works it cites.
Self-normalizing neural networks
Klambauer, Günter, Unterthiner, Thomas, Mayr, Andreas, and Hochreiter, Sepp · 2017
Later among the works it cites.
On the asymptotic optimality of finite approximations to markov decision processes with Borel spaces
Saldi, Naci, Yüksel, Serdar, and Linder, Tamás · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
Silver, David, Schrittwieser, Julian, Simonyan, Karen, Antonoglou, Ioannis, Huang, Aja, Guez, Arthur, Hubert, Thomas, Baker, Lucas, Lai, Matthew, Bolton, Adrian, et al · 2017
Later among the works it cites.