Fetching the paper…
Reading the bibliography…
In this work, we consider the popular tree-based search strategy within the framework of reinforcement learning, the Monte Carlo Tree Search (MCTS), in the context of infinite-horizon discounted cost Markov Decision Process (MDP).
Audibert JY, Munos R, Szepesvári C (2009) Exploration–exploitation tradeoff using variance estimates in multi-armed bandits. Theoretical Computer Science 410(19):1876–1902
1902
Earlier work this paper cites.
Mnih V, Badia AP, Mirza M, Graves A, Lillicrap T, Harley T, Silver D, Kavukcuoglu K (2016) Asynchronous methods for deep reinforcement learning. International conference on machine learning , 1928–1937
1937
Earlier work this paper cites.
Hoeffding W (1963) Probability inequalities for sums of bounded random variables. Journal of Americal Statistics Association 58
1963
Earlier work this paper cites.
Bertsekas D (1975) Convergence of discretization procedures in dynamic programming. IEEE Transactions on Automatic Control 20(3):415–419
1975
Earlier work this paper cites.
Stone CJ (1982) Optimal global rates of convergence for nonparametric regression. The Annals of Statistics 1040–1053
1982
Earlier work this paper cites.
Sutton RS (1988) Learning to predict by the methods of temporal differences. Machine learning 3(1):9–44
1988
Earlier work this paper cites.
Watkins CJ, Dayan P (1992) Q-learning. Machine learning 8(3-4):279–292
1992
Earlier work this paper cites.
Agrawal R (1995) Sample mean based index policies by o (log n) regret for the multi-armed bandit problem. Advances in Applied Probability 27(4):1054–1078
1995
Earlier work this paper cites.
Auer P, Cesa-Bianchi N, Fischer P (2002) Finite-time analysis of the multiarmed bandit problem. Machine learning 47(2-3):235–256
2002
Earlier work this paper cites.
Kearns M, Mansour Y, Ng AY (2002) A sparse sampling algorithm for near-optimal planning in large markov decision processes. Machine learning 49(2-3):193–208
2002
Earlier work this paper cites.
Even-Dar E, Mansour Y (2004) Learning rates for Q-learning. JMLR 5
2004
Earlier work this paper cites.
Chang HS, Fu MC, Hu J, Marcus SI (2005) An adaptive sampling algorithm for solving markov decision processes. Operations Research 53(1):126–139
2005
Earlier work this paper cites.
Coulom R (2006) Efficient selectivity and backup operators in monte-carlo tree search. International conference on computers and games , 72–83 (Springer)
2006
Earlier work this paper cites.
Kocsis L, Szepesvári C (2006) Bandit based monte-carlo planning. European conference on machine learning , 282–293 (Springer)
2006
Earlier work this paper cites.
Kocsis L, Szepesvári C, Willemson J (2006) Improved monte-carlo search. Univ. Tartu, Estonia, Tech. Rep
2006
Earlier work this paper cites.
Strehl AL, Li L, Wiewiora E, Langford J, Littman ML (2006) Pac model-free reinforcement learning. Proceedings of the 23rd international conference on Machine learning , 881–888 (ACM)
2006
Cited alongside, same era.
Coquelin PA, Munos R (2007) Bandit algorithms for tree search. arXiv preprint cs/0703062
2007
Cited alongside, same era.
Hren JF, Munos R (2008) Optimistic planning of deterministic systems. European Workshop on Reinforcement Learning , 151–164 (Springer)
2008
Cited alongside, same era.
Schadd MPD, Winands MHM, van den Herik HJ, Chaslot GMJB, Uiterwijk JWHM (2008) Single-player monte-carlo tree search. van den Herik HJ, Xu X, Ma Z, Winands MHM, eds., Computers and Games , 1–12 (Berlin, Heidelberg: Springer Berlin Heidelberg)
2008
Cited alongside, same era.
Sturtevant NR (2008) An analysis of uct in multi-player games. van den Herik HJ, Xu X, Ma Z, Winands MHM, eds., Computers and Games , 37–49 (Berlin, Heidelberg: Springer Berlin Heidelberg)
Mnih V, Kavukcuoglu K, Silver D, Rusu AA, Veness J, Bellemare MG, Graves A, Riedmiller M, Fidjeland AK, Ostrovski G, et al. (2015) Human-level control through deep reinforcement learning. Nature 518(7540):529
2015
Later among the works it cites.
Schulman J, Levine S, Abbeel P, Jordan M, Moritz P (2015) Trust region policy optimization. International Conference on Machine Learning , 1889–1897
2015
Later among the works it cites.
Silver D, Huang A, Maddison CJ, Guez A, Sifre L, Van Den Driessche G, Schrittwieser J, Antonoglou I, Panneershelvam V, Lanctot M, et al. (2016) Mastering the game of go with deep neural networks and tree search. nature 529(7587):484–489
2016
Later among the works it cites.
Van Hasselt H, Guez A, Silver D (2016) Deep reinforcement learning with double q-learning. AAAI , volume 2, 5 (Phoenix, AZ)
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2008
Cited alongside, same era.
Tsybakov AB (2009) Introduction to Nonparametric Estimation . Springer Series in Statistics (Springer)
2009
Cited alongside, same era.
Salomon A, Audibert JY (2011) Deviations of stochastic bandit regret. International Conference on Algorithmic Learning Theory , 159–173 (Springer)
2011
Cited alongside, same era.
Browne CB, Powley E, Whitehouse D, Lucas SM, Cowling PI, Rohlfshagen P, Tavener S, Perez D, Samothrakis S, Colton S (2012) A survey of monte carlo tree search methods. IEEE Transactions on Computational Intelligence and AI in games 4(1):1–43
2012
Cited alongside, same era.
Dufour F, Prieto-Rumeau T (2012) Approximation of Markov decision processes with general state space. Journal of Mathematical Analysis and applications 388(2):1254–1267
2012
Cited alongside, same era.
Dufour F, Prieto-Rumeau T (2013) Finite linear programming approximations of constrained discounted Markov decision processes. SIAM Journal on Control and Optimization 51(2):1298–1324
2013
Cited alongside, same era.
2013
Cited alongside, same era.
Guo X, Singh S, Lee H, Lewis RL, Wang X (2014) Deep learning for real-time atari game play using offline monte-carlo tree search planning. Advances in neural information processing systems , 3338–3346
2014
Cited alongside, same era.
2017
Later among the works it cites.
Kaufmann E, Koolen WM (2017) Monte-carlo tree search by best arm identification. Advances in Neural Information Processing Systems , 4897–4906
2017
Later among the works it cites.
2017
Later among the works it cites.
2018
Later among the works it cites.
Efroni Y, Dalal G, Scherrer B, Mannor S (2018) Multiple-step greedy policies in approximate and online reinforcement learning. Advances in Neural Information Processing Systems , 5244–5253
2018
Later among the works it cites.
Jiang DR, Ekwedike E, Liu H (2018) Feedback-based tree search for reinforcement learning. International conference on machine learning
2018
Later among the works it cites.
Shah D, Xie Q (2018) Q-learning with nearest neighbors. Bengio S, Wallach H, Larochelle H, Grauman K, Cesa-Bianchi N, Garnett R, eds., Advances in Neural Information Processing Systems 31 , 3115–3125 (Curran Associates, Inc.)
2018
Later among the works it cites.
2018
Later among the works it cites.
Bartlett P, Gabillon V, Healey J, Valko M (2019) Scale-free adaptive planning for deterministic dynamics & discounted rewards. International Conference on Machine Learning , 495–504
2019
Closest in time.
Szepesvári C (2019) Personal communication
2019
Closest in time.