Fetching the paper…
Reading the bibliography…
Reinforcement Learning is divided in two main paradigms: model-free and model-based.
A new method of locating the maximum point of an arbitrary multipeak curve in the presence of noise
H. J. Kushner · 1964
Earlier work this paper cites.
On Bayesian methods for seeking the extremum
J. Močkus · 1975
Earlier work this paper cites.
Models of bounded rationality: Empirically grounded economic reason , volume 3
H. A. Simon · 1982
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
R. S. Sutton · 1991
Earlier work this paper cites.
Modifications of the DIRECT algorithm
J. M. Gablonsky et al · 2001
Earlier work this paper cites.
Gaussian process model based predictive control
J. Kocijan, R. Murray-Smith, C. E. Rasmussen, and A. Girard · 2004
Earlier work this paper cites.
Gaussian processes for machine learning
C. E. Rasmussen and C. K. I. Williams · 2006
Earlier work this paper cites.
Model learning with local Gaussian process regression
D. Nguyen-Tuong, M. Seeger, and J. Peters · 2009
Earlier work this paper cites.
Gaussian processes for global optimization
M. A. Osborne, R. Garnett, and S. J. Roberts · 2009
Earlier work this paper cites.
States versus rewards: dissociable neural prediction error signals underlying model-based and model-free reinforcement learning
J. Gläscher, N. Daw, P. Dayan, and J. P. O’Doherty · 2010
Earlier work this paper cites.
Model learning for robot control: a survey
D. Nguyen-Tuong and J. Peters · 2011
Earlier work this paper cites.
PILCO: A model-based and data-efficient approach to policy search
M. Deisenroth and C. E. Rasmussen · 2011
Cited alongside, same era.
GPy: A Gaussian process framework in python
GPy · 2012
Cited alongside, same era.
MuJoCo: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Cited alongside, same era.
A survey on policy search for robotics
M. P. Deisenroth, G. Neumann, and J. Peters · 2013
Cited alongside, same era.
Learning neural network policies with guided policy search under unknown dynamics
S. Levine and P. Abbeel · 2014
Cited alongside, same era.
Probabilistic differential dynamic programming
Y. Pan and E. Theodorou · 2014
Cited alongside, same era.
Sparse Gaussian process regression for compliant, real-time robot control
J. Schreiter, P. Englert, D. Nguyen-Tuong, and M. Toussaint · 2015
Later among the works it cites.
Bayesian optimization for learning gaits under uncertainty
R. Calandra, A. Seyfarth, J. Peters, and M. P. Deisenroth · 2015
Later among the works it cites.
Learning continuous control policies by stochastic value gradients
N. Heess, G. Wayne, D. Silver, T. Lillicrap, T. Erez, and Y. Tassa · 2015
Later among the works it cites.
Continuous deep Q-learning with model-based acceleration
S. Gu, T. Lillicrap, I. Sutskever, and S. Levine · 2016
Later among the works it cites.
Taking the human out of the loop: A review of Bayesian optimization
B. Shahriari, K. Swersky, Z. Wang, R. P. Adams, and N. de Freitas · 2016
Later among the works it cites.
OpenAI Gym, 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nonlinear modelling and control using Gaussian processes
A. McHutchon · 2014
Cited alongside, same era.
Learning of closed-loop motion control
F. Farshidian, M. Neunert, and J. Buchli · 2014
Cited alongside, same era.
Using trajectory data to improve bayesian optimization for reinforcement learning
A. Wilson, A. Fern, and P. Tadepalli · 2014
Cited alongside, same era.
Model-based contextual policy search for data-efficient generalization of robot skills
A. Kupcsik, M. P. Deisenroth, J. Peters, A. P. Loh, P. Vadakkepat, and G. Neumann · 2014
Cited alongside, same era.
Gaussian processes for data-efficient learning in robotics and control
M. P. Deisenroth, D. Fox, and C. E. Rasmussen · 2015
Cited alongside, same era.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Later among the works it cites.
Goal-driven dynamics learning via Bayesian optimization
S. Bansal, R. Calandra, T. Xiao, S. Levine, and C. J. Tomlin · 2017
Closest in time.
Task-based end-to-end model learning
P. L. Donti, B. Amos, and J. Z. Kolter · 2017
Closest in time.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
A. Nagabandi, G. Kahn, R. S. Fearing, and S. Levine · 2017
Closest in time.
Combining model-based and model-free updates for trajectory-centric reinforcement learning
Y. Chebotar, K. Hausman, M. Zhang, G. Sukhatme, S. Schaal, and S. Levine · 2017
Closest in time.