Fetching the paper…
Reading the bibliography…
Model-based planning holds great promise for improving both sample efficiency and generalization in reinforcement learning (RL).
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
R. S. Sutton · 1990
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
G. E. Hinton, S. Osindero, and Y.-W. Teh · 2006
Earlier work this paper cites.
What is intrinsic motivation? a typology of computational approaches
P.-Y. Oudeyer and F. Kaplan · 2009
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
J. Schmidhuber · 2010
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
M. Deisenroth and C. E. Rasmussen · 2011
Earlier work this paper cites.
Stomp: Stochastic trajectory optimization for motion planning
M. Kalakrishnan, S. Chitta, E. Theodorou, P. Pastor, and S. Schaal · 2011
Earlier work this paper cites.
Discovery of complex behaviors through contact-invariant optimization
I. Mordatch, E. Todorov, and Z. Popović · 2012
Earlier work this paper cites.
Trajectory optimization for domains with contacts using inverse dynamics
T. Erez and E. Todorov · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Earlier work this paper cites.
Model-less feedback control of continuum manipulators in constrained environments
M. C. Yip and D. B. Camarillo · 2014
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Earlier work this paper cites.
Vizdoom: A doom-based ai research platform for visual reinforcement learning
M. Kempka, M. Wydmuch, G. Runc, J. Toczek, and W. Jaśkowski · 2016
Cited alongside, same era.
Combining model-based policy search with online model learning for control of physical humanoids
I. Mordatch, N. Mishra, C. Eppner, and P. Abbeel · 2016
Cited alongside, same era.
Deep directed generative models with energy-based probability estimation
T. Kim and Y. Bengio · 2016
Cited alongside, same era.
Vime: Variational information maximizing exploration
R. Houthooft, X. Chen, Y. Duan, J. Schulman, F. De Turck, and P. Abbeel · 2016
Cited alongside, same era.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Cited alongside, same era.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Later among the works it cites.
P. Dhariwal, C. Hesse, O. Klimov, A. Nichol, M. Plappert, A. Radford, J. Schulman, S. Sidor, Y. Wu, and P. Zhokhov · 2017
Later among the works it cites.
Openai five, 2018
S. Zhang, M. Petrov, P. Jacob, H. Pondé, B. Chan, F. Wolski, S. Sidor, R. Józefowicz, P. Dębiak, D. Farhi, G. Brockman, J. Raiman, J. Tang, C. Dennison, P. Christiano, S. Hashme, L. Schiavo, I. Sutskever, E. Sigler, J. Schneider, J. Schulman, C. Hesse, J. Clark, Q. Fischer, D. Yoon, C. Berner, S. Gray, A. Radford, and D. Luan · 2018
Later among the works it cites.
Learning dexterous in-hand manipulation
M. Andrychowicz, B. Baker, M. Chociej, R. Jozefowicz, B. McGrew, J. Pachocki, A. Petron, M. Plappert, G. Powell, A. Ray, et al · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Prediction and control with temporal segment models
N. Mishra, P. Abbeel, and I. Mordatch · 2017
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
T. Haarnoja, H. Tang, P. Abbeel, and S. Levine · 2017
Cited alongside, same era.
Model predictive path integral control: From theory to parallel computation
G. Williams, A. Aldrich, and E. A. Theodorou · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Cited alongside, same era.
# exploration: A study of count-based exploration for deep reinforcement learning
H. Tang, R. Houthooft, D. Foote, A. Stooke, O. X. Chen, Y. Duan, J. Schulman, F. DeTurck, and P. Abbeel · 2017
Cited alongside, same era.
Information theoretic mpc for model-based reinforcement learning
G. Williams, N. Wagener, B. Goldfain, P. Drews, J. M. Rehg, B. Boots, and E. A. Theodorou · 2017
Cited alongside, same era.
Improving pilco with bayesian neural network dynamics models
Y. Gal, R. McAllister, and C. E. Rasmussen
Cited in the paper.
A. Nichol, V. Pfau, C. Hesse, O. Klimov, and J. Schulman · 2018
Later among the works it cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
S. Levine · 2018
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
A. Nagabandi, G. Kahn, R. S. Fearing, and S. Levine · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Later among the works it cites.
Task-agnostic dynamics priors for deep reinforcement learning
Y. Du and K. Narasimhan · 2019
Closest in time.
Implicit generation and generalization in energy-based models
Y. Du and I. Mordatch · 2019
Closest in time.
Exponential family estimation via adversarial dynamics embedding
B. Dai, Z. Liu, H. Dai, N. He, A. Gretton, L. Song, and D. Schuurmans · 2019
Closest in time.