Fetching the paper…
Reading the bibliography…
We present a new model-based algorithm for reinforcement learning (RL) which consists of explicit exploration and exploitation phases, and is applicable in large or infinite state spaces.
Efficient memory-based learning for robot control
A. W. Moore · 1990
Earlier work this paper cites.
Mixture density networks
C. M. Bishop · 1994
Earlier work this paper cites.
Improving generalization with active learning
D. Cohn, L. Atlas, and R. Ladner · 1994
Earlier work this paper cites.
Generalization in reinforcement learning: Successful examples using sparse coarse coding
R. S. Sutton · 1996
Earlier work this paper cites.
A comparison of direct and model-based reinforcement learning
C. G. Atkeson and J. C. Santamaria · 1997
Earlier work this paper cites.
Integral probability metrics and their generating classes of functions
A. Mueller · 1997
Earlier work this paper cites.
Efficient reinforcement learning in factored mdps
M. Kearns and D. Koller · 1999
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
M. Kearns and S. Singh · 2002
Earlier work this paper cites.
R-max - a general polynomial time algorithm for near-optimal reinforcement learning
R. I. Brafman and M. Tennenholtz · 2003
Earlier work this paper cites.
Exploration in metric state spaces
S. M. Kakade, M. Kearns, and J. Langford · 2003
Earlier work this paper cites.
A theoretical analysis of model-based interval estimation
A. L. Strehl and M. L. Littman · 2005
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
R. Coulom · 2006
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
A. L. Strehl and M. L. Littman · 2008
Earlier work this paper cites.
Dyna-style planning with linear function approximation and prioritized sweeping
R. S. Sutton, C. Szepesvári, A. Geramifard, and M. Bowling · 2008
Earlier work this paper cites.
Near-bayesian exploration in polynomial time
J. Z. Kolter and A. Y. Ng · 2009
Earlier work this paper cites.
Variance-based rewards for approximate bayesian reinforcement learning
J. Sorg, S. Singh, and R. L. Lewis · 2010
Earlier work this paper cites.
Regret bounds for the adaptive control of linear quadratic systems
Y. Abbasi-Yadkori and C. Szepesvári · 2011
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
M. P. Deisenroth and C. E. Rasmussen · 2011
Cited alongside, same era.
Auto-encoding variational bayes, 2013
D. P. Kingma and M. Welling · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization, 2014
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
M. G. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos · 2016
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Later among the works it cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Later among the works it cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
K. Chua, R. Calandra, R. McAllister, and S. Levine · 2018
Later among the works it cites.
Practical contextual bandits with regression oracles
D. J. Foster, A. Agarwal, M. Dudík, H. Luo, and R. E. Schapire · 2018
Later among the works it cites.
Recurrent world models facilitate policy evolution
D. Ha and J. Schmidhuber · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Openai gym, 2016
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
H. v. Hasselt, A. Guez, and D. Silver · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Deep exploration via bootstrapped dqn
I. Osband, C. Blundell, A. Pritzel, and B. Van Roy · 2016
Cited alongside, same era.
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2016
Cited alongside, same era.
On the sample complexity of the linear quadratic regulator
S. Dean, H. Mania, N. Matni, B. Recht, and S. Tu · 2017
Cited alongside, same era.
D. Hafner, T. P. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson · 2018
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
M. Hessel, J. Modayil, H. P. van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. G. Azar, and D. Silver · 2018
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
A. Nagabandi, G. Kahn, R. S. Fearing, and S. Levine · 2018
Later among the works it cites.
Modularized implementation of deep rl algorithms in pytorch
Z. Shangtong · 2018
Later among the works it cites.
Model-based active exploration
P. Shyam, W. Jaskowski, and F. Gomez · 2018
Later among the works it cites.
Universal planning networks: Learning generalizable representations for visuomotor control
A. Srinivas, A. Jabri, P. Abbeel, S. Levine, and C. Finn · 2018
Later among the works it cites.
Model-based reinforcement learning in contextual decision processes
W. Sun, N. Jiang, A. Krishnamurthy, A. Agarwal, and J. Langford · 2018
Later among the works it cites.
Exploration by random network distillation
Y. Burda, H. Edwards, A. Storkey, and O. Klimov · 2019
Closest in time.
Provably efficient RL with rich observations via latent state decoding
S. Du, A. Krishnamurthy, N. Jiang, A. Agarwal, M. Dudik, and J. Langford · 2019
Closest in time.
Model-predictive policy learning with uncertainty regularization for driving in dense traffic
M. Henaff, A. Canziani, and Y. LeCun · 2019
Closest in time.
Algorithmic framework for model-based deep reinforcement learning with theoretical guarantees
Y. Luo, H. Xu, Y. Li, Y. Tian, T. Darrell, and T. Ma · 2019
Closest in time.
Self-supervised exploration via disagreement
D. Pathak, D. Gandhi, and A. Gupta · 2019
Closest in time.