State of the art—a survey of partially observable markov decision processes: theory, models, and algorithms
George E Monahan · 1982
Earlier work this paper cites.
On the gittins index for multiarmed bandits
R. Weber · 1992
Earlier work this paper cites.
Learning without state-estimation in partially observable markovian decision processes
Satinder P. Singh, Tommi S. Jaakkola, and Michael I. Jordan · 1994
Earlier work this paper cites.
Bayesian q-learning
R. Dearden, N. Friedman, and Stuart J. Russell · 1998
Earlier work this paper cites.
Model based bayesian exploration
R. Dearden, N. Friedman, and D. Andre · 1999
Earlier work this paper cites.
A bayesian framework for reinforcement learning
M. Strens · 2000
Earlier work this paper cites.
Optimal learning: computational procedures for bayes-adaptive markov decision processes
M. Duff and A. Barto · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
An analytic solution to discrete bayesian reinforcement learning
P. Poupart, N. Vlassis, J. Hoey, and K. Regan · 2006
Earlier work this paper cites.
Bayesian inverse reinforcement learning
Deepak Ramachandran and E. Amir · 2007
Earlier work this paper cites.
Bayes-adaptive pomdps
Stephane Ross, Brahim Chaib-draa, and Joelle Pineau · 2007
Earlier work this paper cites.
Bayesian multi-task reinforcement learning
A. Lazaric and M. Ghavamzadeh · 2010
Earlier work this paper cites.
Protecting against evaluation overfitting in empirical reinforcement learning
S. Whiteson, B. Tanner, Matthew E. Taylor, and P. Stone · 2011
Earlier work this paper cites.
Learning to grasp under uncertainty
F. Stulp, E. Theodorou, J. Buchli, and S. Schaal · 2011
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Earlier work this paper cites.
Learning to optimize via posterior sampling
Daniel Russo and Benjamin Van Roy · 2014
Earlier work this paper cites.
Bayesian reinforcement learning: A survey
M. Ghavamzadeh, Shie Mannor, Joelle Pineau, and Aviv Tamar · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Earlier work this paper cites.
Geometry and determinism of optimal stationary control in partially observable markov decision processes
Original
Guido Montúfar, K. Zahedi, and N. Ay · 2015
Earlier work this paper cites.
Deep exploration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Earlier work this paper cites.