Fetching the paper…
Reading the bibliography…
Model-based reinforcement learning (RL) algorithms allow us to combine model-generated data with those collected from interaction with the real system in order to alleviate the data efficiency problem in RL.
Dynamic programming and optimal control , volume 2
Dimitri P Bertsekas, Dimitri P Bertsekas, Dimitri P Bertsekas, and Dimitri P Bertsekas · 1995
Earlier work this paper cites.
Using expectation-maximization for reinforcement learning
P. Dayan and G. Hinton · 1997
Earlier work this paper cites.
Q-learning for risk-sensitive control
Vivek S Borkar · 2002
Earlier work this paper cites.
Reinforcement learning by reward-weighted regression for operational space control
J. Peters and S. Schal · 2007
Earlier work this paper cites.
Estimating divergence functionals and the likelihood ratio by penalized convex risk minimization
XuanLong Nguyen, Martin J Wainwright, and Michael I Jordan · 2008
Earlier work this paper cites.
General duality between optimal control and estimation
E. Todorov · 2008
Earlier work this paper cites.
Efficient sample reuse in EM-based policy search
Peters J. Hachiya, H. and M. Sugiyama · 2009
Earlier work this paper cites.
Robot trajectory optimization using approximate inference
M. Toussaint · 2009
Earlier work this paper cites.
Relative entropy policy search
J. Peters, K. Mulling, and Y. Altun · 2010
Earlier work this paper cites.
Variational inference for policy search in changing situations
G. Neumann · 2011
Earlier work this paper cites.
Optimal control as a graphical model inference problem
H. Kappen, V. Gomez, and M. Opper · 2012
Earlier work this paper cites.
Stochastic variational inference
M. Hoffman, D. Blei, C. Wang, and J. Paisley · 2013
Earlier work this paper cites.
Variational policy search via trajectory optimization
S. Levine and V. Koltun · 2013
Cited alongside, same era.
Playing Atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Cited alongside, same era.
On stochastic optimal control and reinforcement learning by approximate inference
Toussaint M. Rawlik, K. and S. Vijayakumar · 2013
Cited alongside, same era.
Learning neural network policies with guided policy search under unknown dynamics
S. Levine and P. Abbeel · 2014
Cited alongside, same era.
Learning complex neural network policies with trajectory optimization
S. Levine and V. Koltun · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Jacob Buckman, Danijar Hafner, George Tucker, Eugene Brevdo, and Honglak Lee · 2018
Later among the works it cites.
Unsupervised exploration with deep model-based reinforcement learning
Kurtland Chua, Rowan McAllister, Roberto Calandra, and Sergey Levine · 2018
Later among the works it cites.
Iterative value-aware model learning
A. Farahmand · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Later among the works it cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
S. Levine · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. Mnih, K. Kavukcuoglu, D. Silver, A. Rusu, J. Veness, M. Bellemare, A. Graves, M. Riedmiller, A. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Cited alongside, same era.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Cited alongside, same era.
Embed to control: A locally linear latent dynamics model for control from raw images
M. Watter, J. Springenberg, J. Boedecker, and M. Riedmiller · 2015
Cited alongside, same era.
Path integral guided policy search
Y. Chebotar, M. Kalakrishnan, A. Yahya, A. Li, S. Schaal, and S. Levine · 2017
Cited alongside, same era.
Bridging the gap between value and policy based reinforcement learning
O. Nachum, M. Norouzi, K. Xu, and D. Schuurmans · 2017
Cited alongside, same era.
Maximum a posteriori policy optimisation
A. Abdolmaleki, J. Springenberg, Y. Tassa, R. Munos, N. Heess, and M. Riedmiller · 2018
Cited alongside, same era.
O. Nachum, Y. Chow, and M. Ghavamzadeh · 2018
Later among the works it cites.
Reinforcement learning: An introduction
R. Sutton and A. Barto · 2018
Later among the works it cites.
VIREL: A variational inference framework for reinforcement learning
M. Fellows, A. Mahajan, T. Rudner, and S. Whiteson · 2019
Later among the works it cites.
When to trust your model: Model-based policy optimization
M. Janner, J. Fu, M. Zhang, and S. Levine · 2019
Later among the works it cites.
Policy-aware model learning for policy gradient methods
R. Abachi, M. Ghavamzadeh, and A. Farahmand · 2020
Closest in time.
Prediction, consistency, curvature: Representation learning for locally-linear control
N. Levine, Y. Chow, R. Shu, A. Li, M. Ghavamzadeh, and H. Bui · 2020
Closest in time.