Making the world differentiable: On using self-supervised fully recurrent neural networks for dynamic reinforcement learning and planning in non-stationary environments
J. Schmidhuber · 1990
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
L.-J. Lin · 1992
Earlier work this paper cites.
Learning to drive a bicycle using reinforcement learning and shaping
J. Randløv and P. Alstrøm · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
A. Y. Ng, D. Harada, and S. Russell · 1999
Earlier work this paper cites.
Fastslam: A factored solution to the simultaneous localization and mapping problem
M. Montemerlo, S. Thrun, D. Koller, B. Wegbreit, et al · 2002
Earlier work this paper cites.
Design and use paradigms for gazebo, an open-source multi-robot simulator
N. P. Koenig and A. Howard · 2004
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
P. Vincent, H. Larochelle, Y. Bengio, and P.-A. Manzagol · 2008
Earlier work this paper cites.
Rep: 119-specification for turtlebot compatible platforms, dec. 2011
M. Wise and T. Foote · 2011
Earlier work this paper cites.
Auto-encoding variational bayes
Original
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Original
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
Universal value function approximators
T. Schaul, D. Horgan, K. Gregor, and D. Silver · 2015
Earlier work this paper cites.