Lqr-trees: Feedback motion planning via sums-of-squares verification
R. Tedrake, I. R. Manchester, M. Tobenkin, and J. W. Roberts · 2010
Cited alongside, same era.
Curriculum learning for motor skills
A. Karpathy and M. Van De Panne · 2012
Cited alongside, same era.
Continually adding self-invented problems to the repertoire: First experiments with POWERPLAY
R. K. Srivastava, B. R. Steunebrink, M. Stollenga, and J. Schmidhuber · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Cited alongside, same era.
POWER P LAY : Training an Increasingly General Problem Solver by Continually Searching for the Simplest Still Unsolvable Problem
J. Schmidhuber · 2013
Cited alongside, same era.
Active learning of inverse models with intrinsically motivated goal exploration in robots
A. Baranes and P.-Y. Oudeyer · 2013
Cited alongside, same era.
A survey on policy search for robotics
M. P. Deisenroth, G. Neumann, J. Peters, et al · 2013
Cited alongside, same era.
Learning to execute
Original
W. Zaremba and I. Sutskever · 2014
Cited alongside, same era.
High-Dimensional continuous control using generalized advantage estimation
J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Original
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Cited alongside, same era.
Scheduled sampling for sequence prediction with recurrent neural networks
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer · 2015
Cited alongside, same era.