Fetching the paper…
Reading the bibliography…
A major challenge in modern reinforcement learning (RL) is efficient control of dynamical systems from high-dimensional sensory observations.
Swing up control of inverted pendulum
K. Furuta, M. Yamakita, and S. Kobayashi · 1991
Earlier work this paper cites.
A cartpole experiment benchmark for trainable controllers
S. Geva and J. Sitte · 1993
Earlier work this paper cites.
The swing up control problem for the acrobot
M. Spong · 1995
Earlier work this paper cites.
Q-learning for risk-sensitive control
V. Borkar · 2002
Earlier work this paper cites.
Iterative linear quadratic regulator design for nonlinear biological movement systems
W. Li and E. Todorov · 2004
Earlier work this paper cites.
Principles of guidance-based path following in 2D and 3D
M. Breivik and T. Fossen · 2005
Earlier work this paper cites.
Optimization of convex risk functions
A. Ruszczyński and A. Shapiro · 2006
Earlier work this paper cites.
PILCO: A model-based and data-efficient approach to policy search
M. Deisenroth and C. Rasmussen · 2011
Earlier work this paper cites.
Auto-encoding variational bayes, 2013
D. Kingma and M. Welling · 2013
Earlier work this paper cites.
Playing Atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization, 2014
D. Kingma and J. Ba · 2014
Cited alongside, same era.
Singularity-avoiding swing-up control for underactuated three-link gymnast robot using virtual coupling between control torques
X. Lai, A. Zhang, M. Wu, and J. She · 2015
Cited alongside, same era.
Embed to control: A locally linear latent dynamics model for control from raw images
M. Watter, J. Springenberg, J. Boedecker, and M. Riedmiller · 2015
Cited alongside, same era.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Later among the works it cites.
When to trust your model: Model-based policy optimization
M. Janner, J. Fu, M. Zhang, and S. Levine · 2019
Later among the works it cites.
SOLAR: Deep structured representations for model-based reinforcement learning
M. Zhang, S. Vikram, L. Smith, P. Abbeel, M. Johnson, and S. Levine · 2019
Later among the works it cites.
Policy-aware model learning for policy gradient methods
R. Abachi, M. Ghavamzadeh, and A. Farahmand · 2020
Closest in time.
Variational model-based policy optimization
Y. Chow, B. Cui, M. Ryu, and M. Ghavamzadeh · 2020
Closest in time.
Dream to control: Learning behaviors by latent imagination
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Robust locally-linear controllable embedding
E. Banijamali, R. Shu, M. Ghavamzadeh, H. Bui, and A. Ghodsi · 2018
Cited alongside, same era.
Visual Foresight: Model-based deep reinforcement learning for vision-based robotic control
F. Ebert, C. Finn, S. Dasari, A. Xie, A. Lee, and S. Levine · 2018
Cited alongside, same era.
Iterative value-aware model learning
A. Farahmand · 2018
Cited alongside, same era.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz
Cited in the paper.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz
Cited in the paper.
D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi · 2020
Closest in time.
Prediction, consistency, curvature: Representation learning for locally-linear control
N. Levine, Y. Chow, R. Shu, A. Li, M. Ghavamzadeh, and H. Bui · 2020
Closest in time.
Predictive coding for locally-linear control
R. Shu, T. Nguyen, Y. Chow, T. Pham, K. Than, M. Ghavamzadeh, S. Ermon, and H. Bui · 2020
Closest in time.