Fetching the paper…
Reading the bibliography…
We propose MDP-GapE, a new trajectory-based Monte-Carlo Tree Search algorithm for planning in a Markov Decision Process in which transitions have a finite support.
A Sparse Sampling Algorithm for Near-Optimal Planning in Large Markov Decision Processes
Michael J. Kearns, Yishay Mansour, and Andrew Y. Ng · 2002
Earlier work this paper cites.
Self-normalized processes: exponential inequalities, moment bounds and iterated logarithm laws
Victor H de la Pena, Michael J Klass, and Tze Leung Lai · 2004
Earlier work this paper cites.
Bandit Based Monte-Carlo Planning
Levente Kocsis and Csaba Szepesvári · 2006
Earlier work this paper cites.
Optimistic planning of deterministic systems
Jean-Francois Hren and Rémi Munos · 2008
Earlier work this paper cites.
Open loop optimistic planning
S Bubeck and R Munos · 2010
Earlier work this paper cites.
Optimism in Reinforcement Learning and Kullback-Leibler Divergence
S. Filippi, O. Cappé, and A. Garivier · 2010
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
The KL-UCB algorithm for bounded stochastic bandits and beyond
Aurélien Garivier and Olivier Cappé · 2011
Earlier work this paper cites.
A Survey of Monte Carlo Tree Search Methods
C. Browne, E. Powley, D. Whitehouse, S. Lucas, P. Cowling, P. Rohlfshagen, S. Tavener, D. Perez, S. Samothrakis, and S. Colton · 2012
Earlier work this paper cites.
Optimistic planning for Markov decision processes
Lucian Busoniu and Rémi Munos · 2012
Earlier work this paper cites.
Elements of information theory
Thomas M Cover and Joy A Thomas · 2012
Cited alongside, same era.
Best arm identification: A unified approach to fixed budget and fixed confidence
Victor Gabillon, Mohammad Ghavamzadeh, and Alessandro Lazaric · 2012
Cited alongside, same era.
MCTS based on simple regret
David Tolpin and Solomon Eyal Shimony · 2012
Cited alongside, same era.
Kullback-Leibler upper confidence bounds for optimal sequential allocation
O. Cappé, A. Garivier, O-A. Maillard, R. Munos, and G. Stoltz · 2013
Cited alongside, same era.
Simple Regret Optimization in Online Planning for Markov Decision Processes
Zohar Feldman and Carmel Domshlak · 2014
Cited alongside, same era.
From bandits to Monte-Carlo Tree Search: The optimistic principle applied to optimization and planning
R. Munos · 2014
Cited alongside, same era.
Structured Best Arm Identification with Fixed Confidence
Ruitong Huang, Mohammad M. Ajallooeian, Csaba Szepesvári, and Martin Müller · 2017
Later among the works it cites.
Monte-Carlo tree search by best arm identification
Emilie Kaufmann and Wouter M. Koolen · 2017
Later among the works it cites.
Aurélien Garivier, Hédi Hadiji, Pierre Menard, and Gilles Stoltz · 2018
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy P. Lillicrap, Karen Simonyan, and Demis Hassabis · 2018
Later among the works it cites.
Planning in entropy-regularized Markov decision processes and games
Jean-Bastien Grill, Omar Darwiche Domingues, Pierre Ménard, Rémi Munos, and Michal Valko · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Minimizing Simple and Cumulative Regret in Monte-Carlo Tree Search
Tom Pepels, Tristan Cazenave, Mark H. M. Winands, and Marc Lanctot · 2014
Cited alongside, same era.
Optimistic Planning in Markov Decision Processes using a generative model
B. Szorenyi, G. Kedenburg, and R. Munos · 2014
Cited alongside, same era.
Blazing the trails before beating the path: Sample-efficient Monte-Carlo planning
J.-B. Grill, M. Valko, and R. Munos · 2016
Cited alongside, same era.
Unifying PAC and regret: Uniform PAC bounds for episodic reinforcement learning
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Cited alongside, same era.
Practical open-loop optimistic planning
Edouard Leurent and Odalric-Ambrym Maillard · 2019
Later among the works it cites.
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, Timothy P. Lillicrap, and David Silver · 2019
Later among the works it cites.
Non-Asymptotic Gap-Dependent Regret Bounds for Tabular MDPs
Max Simchowitz and Kevin G Jamieson · 2019
Later among the works it cites.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Andrea Zanette and Emma Brunskill · 2019
Later among the works it cites.