Fetching the paper…
Reading the bibliography…
Real Time Dynamic Programming (RTDP) is an online algorithm based on Dynamic Programming (DP) that acts by 1-step greedy planning.
Learning to act using real-time dynamic programming
Andrew Barto, Steven Bradtke, and Satinder Singh · 1995
Earlier work this paper cites.
Neuro-dynamic programming
D. Bertsekas and J. Tsitsiklis · 1996
Earlier work this paper cites.
Model reduction techniques for computing approximately optimal solutions for Markov decision processes
Thomas Dean, Robert Givan, and Sonia Leach · 1997
Earlier work this paper cites.
Abstraction and approximate decision-theoretic planning
Richard Dearden and Craig Boutilier · 1997
Earlier work this paper cites.
Planning with incomplete information as heuristic search in belief space
Blai Bonet and Hector Geffner · 2000
Earlier work this paper cites.
A sparse sampling algorithm for near-optimal planning in large Markov decision processes
Michael Kearns, Yishay Mansour, and Andrew Ng · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
Labeled rtdp: Improving the convergence of real-time dynamic programming
Blai Bonet and Hector Geffner · 2003
Earlier work this paper cites.
Approximate equivalence of markov decision processes
Eyal Even-Dar and Yishay Mansour · 2003
Earlier work this paper cites.
Bounded real-time dynamic programming: Rtdp with monotone upper bounds and performance guarantees
Brendan McMahan, Maxim Likhachev, and Geoffrey Gordon · 2005
Earlier work this paper cites.
Learning in real-time search: A unifying framework
Vadim Bulitko and Greg Lee · 2006
Cited alongside, same era.
Bandit based Monte-Carlo planning
Levente Kocsis and Csaba Szepesvári · 2006
Cited alongside, same era.
Towards a unified theory of state abstraction for MDPs
L. Li, T. Walsh, and M. Littman · 2006
Cited alongside, same era.
PAC reinforcement learning bounds for RTDP and rand-RTDP
A. Strehl, L. Li, and M. Littman · 2006
Cited alongside, same era.
Bandit algorithms for tree search
Pierre-Arnaud Coquelin and Rémi Munos · 2007
Cited alongside, same era.
Performance bounds in l_p-norm for approximate value iteration
Rémi Munos · 2007
Cited alongside, same era.
Algorithmic survey of parametric value function approximation
Matthieu Geist and Olivier Pietquin · 2013
Later among the works it cites.
From bandits to Monte-Carlo tree search: The optimistic principle applied to optimization and planning
Rémi Munos · 2014
Later among the works it cites.
Near optimal behavior via approximate state abstraction
David Abel, D. Hershkowitz, and Michael Littman · 2016
Later among the works it cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Later among the works it cites.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Later among the works it cites.
Multiple-step greedy policies in approximate and online reinforcement learning
Yonathan Efroni, Gal Dalal, Bruno Scherrer, and Shie Mannor · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alexander Strehl, Lihong Li, and Michael Littman · 2009
Cited alongside, same era.
A survey of Monte Carlo tree search methods
Cameron Browne, Edward Powley, Daniel Whitehouse, Simon Lucas, Peter Cowling, Philipp Rohlfshagen, Stephen Tavener, Diego Perez, Spyridon Samothrakis, and Simon Colton · 2012
Cited alongside, same era.
Lrtdp versus uct for online probabilistic planning
Andrey Kolobov, Daniel S Weld, et al · 2012
Cited alongside, same era.
Approximate modified policy iteration
Bruno Scherrer, Mohammad Ghavamzadeh, Victor Gabillon, and Matthieu Geist · 2012
Cited alongside, same era.
Later among the works it cites.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Later among the works it cites.
Variance reduced value iteration and faster algorithms for solving Markov decision processes
Aaron Sidford, Mengdi Wang, Xian Wu, and Yinyu Ye · 2018
Later among the works it cites.
How to combine tree-search methods in reinforcement learning
Y. Efroni, G. Dalal, B. Scherrer, and S. Mannor · 2019
Closest in time.
Tight regret bounds for model-based reinforcement learning with greedy policies
Yonathan Efroni, Nadav Merlis, Mohammad Ghavamzadeh, and Shie Mannor · 2019
Closest in time.