Fetching the paper…
Reading the bibliography…
Deep reinforcement learning (DRL) methods such as the Deep Q-Network (DQN) have achieved state-of-the-art results in a variety of challenging, high-dimensional domains.
Individual comparisons by ranking methods
Wilcoxon, Frank · 1945
Earlier work this paper cites.
Reinforcement learning for robots using neural networks
Lin, Long-Ji · 1993
Earlier work this paper cites.
Improving elevator performance using reinforcement learning
Barto, AG and Crites, RH · 1996
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
Tsitsiklis, John N, Van Roy, Benjamin, et al · 1997
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, Richard and Barto, Andrew · 1998
Earlier work this paper cites.
Least-squares policy iteration
Lagoudakis, Michail G and Parr, Ronald · 2003
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, Damien, Geurts, Pierre, and Wehenkel, Louis · 2005
Earlier work this paper cites.
Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method
Riedmiller, Martin · 2005
Earlier work this paper cites.
Approximate dynamic programming
Bertsekas, Dimitri P · 2008
Earlier work this paper cites.
Regularized policy iteration
Farahmand, Amir M, Ghavamzadeh, Mohammad, Mannor, Shie, and Szepesvári, Csaba · 2009
Earlier work this paper cites.
What is the best multi-stage architecture for object recognition?
Jarrett, Kevin, Kavukcuoglu, Koray, LeCun, Yann, et al · 2009
Earlier work this paper cites.
Regularization and feature selection in least-squares temporal difference learning
Kolter, J Zico and Ng, Andrew Y · 2009
Earlier work this paper cites.
Graying the black box: Understanding dqns
Zahavy, Tom, Ben-Zrihem, Nir, and Mannor, Shie · 2009
Cited alongside, same era.
Bayesian inference in statistical analysis
Box, George EP and Tiao, George C · 2011
Cited alongside, same era.
Approximate dynamic programming via a smoothed linear program
Desai, Vijay V, Farias, Vivek F, and Moallemi, Ciamac C · 2012
Cited alongside, same era.
Neural networks for machine learning lecture 6a overview of mini–batch gradient descent
Hinton, Geoffrey, Srivastava, NiRsh, and Swersky, Kevin · 2012
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Bellemare, Marc G, Naddaf, Yavar, Veness, Joel, and Bowling, Michael · 2013
Cited alongside, same era.
Decaf: A deep convolutional activation feature for generic visual recognition
Donahue, Jeff, Jia, Yangqing, Vinyals, Oriol, Hoffman, Judy, Zhang, Ning, Tzeng, Eric, and Darrell, Trevor · 2013
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, Nitish Shirish, Mudigere, Dheevatsa, Nocedal, Jorge, Smelyanskiy, Mikhail, and Tang, Ping Tak Peter · 2016
Later among the works it cites.
State of the art control of atari games using shallow reinforcement learning
Liang, Yitao, Machado, Marlos C, Talvitie, Erik, and Bowling, Michael · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Mnih, Volodymyr, Badia, Adria Puigdomenech, Mirza, Mehdi, Graves, Alex, Lillicrap, Timothy P, Harley, Tim, Silver, David, and Kavukcuoglu, Koray · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
Silver, David, Huang, Aja, Maddison, Chris J, Guez, Arthur, Sifre, Laurent, Van Den Driessche, George, Schrittwieser, Julian, Antonoglou, Ioannis, Panneershelvam, Veda, Lanctot, Marc, et al · 2016
Later among the works it cites.
Deep reinforcement learning with double q-learning
Van Hasselt, Hado, Guez, Arthur, and Silver, David · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, Diederik and Ba, Jimmy · 2014
Cited alongside, same era.
Bayesian reinforcement learning: A survey
Ghavamzadeh, Mohammad, Mannor, Shie, Pineau, Joelle, Tamar, Aviv, et al · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A, Veness, Joel, Bellemare, Marc G, Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K, Ostrovski, Georg, et al · 2015
Cited alongside, same era.
Approximate modified policy iteration and its application to the game of tetris
Scherrer, Bruno, Ghavamzadeh, Mohammad, Gabillon, Victor, Lesner, Boris, and Geist, Matthieu · 2015
Cited alongside, same era.
Model-free episodic control
Blundell, Charles, Uria, Benigno, Pritzel, Alexander, Li, Yazhe, Ruderman, Avraham, Leibo, Joel Z, Rae, Jack, Wierstra, Daan, and Hassabis, Demis · 2016
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Wang, Ziyu, Schaul, Tom, Hessel, Matteo, van Hasselt, Hado, Lanctot, Marc, and de Freitas, Nando · 2016
Later among the works it cites.
Sharp minima can generalize for deep nets
Dinh, Laurent, Pascanu, Razvan, Bengio, Samy, and Bengio, Yoshua · 2017
Closest in time.
Learning from demonstrations for real world reinforcement learning
Hester, Todd, Vecerik, Matej, Pietquin, Olivier, Lanctot, Marc, Schaul, Tom, Piot, Bilal, Sendonaris, Andrew, Dulac-Arnold, Gabriel, Osband, Ian, Agapiou, John, et al · 2017
Closest in time.
Hoffer, Elad, Hubara, Itay, and Soudry, Daniel · 2017
Closest in time.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, James, Pascanu, Razvan, Rabinowitz, Neil, Veness, Joel, Desjardins, Guillaume, Rusu, Andrei A, Milan, Kieran, Quan, John, Ramalho, Tiago, Grabska-Barwinska, Agnieszka, et al · 2017
Closest in time.
A deep hierarchical approach to lifelong learning in minecraft
Tessler, Chen, Givony, Shahar, Zahavy, Tom, Mankowitz, Daniel J, and Mannor, Shie · 2017
Closest in time.