Fetching the paper…
Reading the bibliography…
State-of-the-art efficient model-based Reinforcement Learning (RL) algorithms typically act by iteratively solving empirical models, i.e., by performing \emph{full-planning} on Markov Decision Processes (MDPs) built by the gathered experience.
Asymptotically efficient adaptive allocation rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Richard S Sutton · 1991
Earlier work this paper cites.
Prioritized sweeping: Reinforcement learning with less data and less time
Andrew W Moore and Christopher G Atkeson · 1993
Earlier work this paper cites.
Learning to act using real-time dynamic programming
Andrew G Barto, Steven J Bradtke, and Satinder P Singh · 1995
Earlier work this paper cites.
Neuro-dynamic programming , volume 5
Dimitri P Bertsekas and John N Tsitsiklis · 1996
Earlier work this paper cites.
Labeled rtdp: Improving the convergence of real-time dynamic programming
Blai Bonet and Hector Geffner · 2003
Earlier work this paper cites.
Inequalities for the l1 deviation of the empirical distribution
Tsachy Weissman, Erik Ordentlich, Gadiel Seroussi, Sergio Verdu, and Marcelo J Weinberger · 2003
Earlier work this paper cites.
Bounded real-time dynamic programming: Rtdp with monotone upper bounds and performance guarantees
H Brendan McMahan, Maxim Likhachev, and Geoffrey J Gordon · 2005
Earlier work this paper cites.
Focused real-time dynamic programming for mdps: Squeezing more out of a heuristic
Trey Smith and Reid Simmons · 2006
Earlier work this paper cites.
Pac reinforcement learning bounds for rtdp and rand-rtdp
Alexander L Strehl, Lihong Li, and Michael L Littman · 2006
Earlier work this paper cites.
Pseudo-maximization and self-normalized processes
Victor H de la Peña, Michael J Klass, Tze Leung Lai, et al · 2007
Earlier work this paper cites.
Self-normalized processes: Limit theory and Statistical Applications
Victor H de la Peña, Tze Leung Lai, and Qi-Man Shao · 2008
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Alexander L Strehl and Michael L Littman · 2008
Cited alongside, same era.
Regal: A regularization based algorithm for reinforcement learning in weakly communicating mdps
Peter L Bartlett and Ambuj Tewari · 2009
Cited alongside, same era.
Empirical bernstein bounds and sample variance penalization
Andreas Maurer and Massimiliano Pontil · 2009
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Cited alongside, same era.
Incremental model-based learners with formal learning-time guarantees
Alexander L Strehl, Lihong Li, and Michael L Littman · 2012
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Later among the works it cites.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Later among the works it cites.
Why is posterior sampling better than optimism for reinforcement learning?
Ian Osband and Benjamin Van Roy · 2017
Later among the works it cites.
Policy certificates: Towards accountable reinforcement learning
Christoph Dann, Lihong Li, Wei Wei, and Emma Brunskill · 2018
Later among the works it cites.
Efficient bias-span-constrained exploration-exploitation in reinforcement learning
Ronan Fruit, Matteo Pirotta, Alessandro Lazaric, and Ronald Ortner · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Cited alongside, same era.
Planning by prioritized sweeping with small backups
Harm Van Seijen and Richard S Sutton · 2013
Cited alongside, same era.
Thompson sampling for learning parameterized markov decision processes
Aditya Gopalan and Shie Mannor · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
On lower bounds for regret in reinforcement learning
Ian Osband and Benjamin Van Roy · 2016
Cited alongside, same era.
Value iteration networks
Aviv Tamar, Yi Wu, Garrett Thomas, Sergey Levine, and Pieter Abbeel · 2016
Cited alongside, same era.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
Shipra Agrawal and Randy Jia · 2017
Cited alongside, same era.
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson · 2018
Later among the works it cites.
Feedback-based tree search for reinforcement learning
Daniel R Jiang, Emmanuel Ekwedike, and Han Liu · 2018
Later among the works it cites.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Later among the works it cites.
Deep dyna-q: Integrating planning for task-completion dialogue policy learning
Baolin Peng, Xiujun Li, Jianfeng Gao, Jingjing Liu, Kam-Fai Wong, and Shang-Yu Su · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Variance-aware regret bounds for undiscounted reinforcement learning in mdps
Mohammad Sadegh Talebi and Odalric-Ambrym Maillard · 2018
Later among the works it cites.
Andrea Zanette and Emma Brunskill · 2019
Closest in time.