Alvinn: An autonomous land vehicle in a neural network
D. A. Pomerleau · 1989
Earlier work this paper cites.
Hybrid reinforcement/supervised learning of dialogue policies from fixed data sets
J. Henderson, O. Lemon, and K. Georgila · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Sample-efficient batch reinforcement learning for dialogue management optimization
O. Pietquin, M. Geist, S. Chandramohan, and H. Frezza-Buet · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Original
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
H. Van Hasselt, A. Guez, and D. Silver · 2016
Earlier work this paper cites.
Coco-text: Dataset and benchmark for text detection and recognition in natural images
Original
A. Veit, T. Matera, L. Neumann, J. Matas, and S. Belongie · 2016
Earlier work this paper cites.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
S. Gu, E. Holly, T. Lillicrap, and S. Levine · 2017
Earlier work this paper cites.
Emergence of locomotion behaviours in rich environments, 2017
N. Heess, D. TB, S. Sriram, J. Lemmon, J. Merel, G. Wayne, Y. Tassa, T. Erez, Z. Wang, S. M. A. Eslami, M. Riedmiller, and D. Silver · 2017
Earlier work this paper cites.
Safe policy improvement with baseline bootstrapping
Original
R. Laroche, P. Trichelair, and R. T. d. Combes · 2017
Earlier work this paper cites.