Fetching the paper…
Reading the bibliography…
Deep reinforcement learning (deep RL) has achieved superior performance in complex sequential tasks by using a deep neural network as its function approximator and by learning directly from raw images.
Asynchronous methods for deep reinforcement learning. In International Conference on Machine Learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. 2016 · 1937
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Long-Ji Lin. 1992 · 1992
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan. 1992 · 1992
Earlier work this paper cites.
Multitask Learning
Rich Caruana. 1998 · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto. 1998 · 1998
Earlier work this paper cites.
A survey of robot learning from demonstration
Brenna D Argall, Sonia Chernova, Manuela Veloso, and Brett Browning. 2009 · 2009
Earlier work this paper cites.
The Difficulty of Training Deep Architectures and the Effect of Unsupervised Pre-Training. In Twelfth International Conference on Artificial Intelligence and Statistics (AISTATS)
Dumitru Erhan, Pierre-Antoine Manzagol, Yoshua Bengio, Samy Bengio, and Pascal Vincent. 2009 · 2009
Earlier work this paper cites.
Learning from imbalanced data
Haibo He and Edwardo A Garcia. 2009 · 2009
Earlier work this paper cites.
Why Does Unsupervised Pre-training Help Deep Learning?
Dumitru Erhan, Yoshua Bengio, Aaron Courville, Pierre-Antoine Manzagol, Pascal Vincent, and Samy Bengio. 2010 · 2010
Cited alongside, same era.
Deep belief nets as function approximators for reinforcement learning
Farnaz Abtahi and Ian Fasel. 2011 · 2011
Cited alongside, same era.
The Arcade Learning Environment: An Evaluation Platform for General Agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling. 2013 · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba. 2014 · 2014
Cited alongside, same era.
How transferable are features in deep neural networks?. In Advances in neural information processing systems
Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson. 2014 · 2014
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Later among the works it cites.
Sample efficient actor-critic with experience replay
Ziyu Wang, Victor Bapst, Nicolas Heess, Volodymyr Mnih, Remi Munos, Koray Kavukcuoglu, and Nando de Freitas. 2016 · 2016
Later among the works it cites.
Deep reinforcement learning from human preferences
Paul Christiano, Jan Leike, Tom B Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Closest in time.
The Atari Grand Challenge Dataset
Vitaly Kurin, Sebastian Nowozin, Katja Hofmann, Lucas Beyer, and Bastian Leibe. 2017 · 2017
Closest in time.
Learning to repeat: Fine grained action repetition for deep reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Faster reinforcement learning after pretraining deep networks to predict state dynamics. In Neural Networks (IJCNN), 2015 International Joint Conference on
Charles W Anderson, Minwoo Lee, and Daniel L Elliott. 2015 · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. 2016 · 2016
Cited alongside, same era.
Sahil Sharma, Aravind S Lakshminarayanan, and Balaraman Ravindran. 2017 · 2017
Closest in time.
StarCraft II: A New Challenge for Reinforcement Learning
Oriol Vinyals, Timo Ewalds, Sergey Bartunov, Petko Georgiev, Alexander Sasha Vezhnevets, Michelle Yeo, Alireza Makhzani, Heinrich Küttler, John Agapiou, Julian Schrittwieser, et al · 2017
Closest in time.
Improving Reinforcement Learning with Confidence-Based Demonstrations. In Proceedings of the 26th International Conference on Artificial Intelligence (IJCAI)
Zhaodong Wang and Matthew E. Taylor. 2017 · 2017
Closest in time.
Deep Q-learning from Demonstrations. In Proceedings of the 32nd AAAI Conference on Artificial Intelligence
Todd Hester, Matej Vecerik, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Andrew Sendonaris, Gabriel Dulac-Arnold, Ian Osband, John Agapiou, Joel Z. Leibo, and Audrunas Gruslys. 2018 · 2018
Closest in time.