Fetching the paper…
Reading the bibliography…
Instability and variability of Deep Reinforcement Learning (DRL) algorithms tend to adversely affect their performance.
A Markovian decision process
Bellman, Richard · 1957
Earlier work this paper cites.
Q-learning
Watkins, Christopher JCH and Dayan, Peter · 1992
Earlier work this paper cites.
Advantage updating
Baird III, Leemon C · 1993
Earlier work this paper cites.
Reinforcement learning for robots using neural networks
Lin, Long-Ji · 1993
Earlier work this paper cites.
Issues in using function approximation for reinforcement learning
Thrun, Sebastian and Schwartz, Anton · 1993
Earlier work this paper cites.
On the convergence of stochastic iterative dynamic programming algorithms
Jaakkola, Tommi, Jordan, Michael I, and Singh, Satinder P · 1994
Earlier work this paper cites.
On-line Q-learning using connectionist systems
Rummery, Gavin A and Niranjan, Mahesan · 1994
Earlier work this paper cites.
Asynchronous stochastic approximation and q-learning
Tsitsiklis, John N · 1994
Earlier work this paper cites.
Generalization in reinforcement learning: Safely approximating the value function
Boyan, Justin and Moore, Andrew W · 1995
Earlier work this paper cites.
Temporal difference learning and td-gammon
Tesauro, Gerald · 1995
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
Tsitsiklis, John N and Van Roy, Benjamin · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Yann, Bottou, Léon, Bengio, Yoshua, and Haffner, Patrick · 1998
Cited alongside, same era.
Reinforcement Learning: An Introduction
Sutton, Richard S and Barto, Andrew G · 1998
Cited alongside, same era.
Policy gradient methods for reinforcement learning with function approximation
Sutton, Richard S, McAllester, David A, Singh, Satinder P, and Mansour, Yishay · 1999
Cited alongside, same era.
Learning rates for q-learning
Even-Dar, Eyal and Mansour, Yishay · 2003
Cited alongside, same era.
Neural fitted Q iteration–first experiences with a data efficient neural reinforcement learning method
Riedmiller, Martin · 2005
Cited alongside, same era.
Double Q-learning
Van Hasselt, Hado · 2010
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A, Veness, Joel, Bellemare, Marc G, Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K, Ostrovski, Georg, et al · 2015
Later among the works it cites.
Schaul, Tom, Quan, John, Antonoglou, Ioannis, and Silver, David · 2015
Later among the works it cites.
Deep reinforcement learning with double Q-learning
Van Hasselt, Hado, Guez, Arthur, and Silver, David · 2015
Later among the works it cites.
Dueling network architectures for deep reinforcement learning
Wang, Ziyu, de Freitas, Nando, and Lanctot, Marc · 2015
Later among the works it cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, Marc G, Srinivasan, Sriram, Ostrovski, Georg, Schaul, Tom, Saxton, David, and Munos, Remi · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Imagenet classification with deep convolutional neural networks
Krizhevsky, Alex, Sutskever, Ilya, and Hinton, Geoffrey E · 2012
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Cited alongside, same era.
Playing Atari with deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Graves, Alex, Antonoglou, Ioannis, Wierstra, Daan, and Riedmiller, Martin · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, Diederik P. and Ba, Jimmy · 2014
Cited alongside, same era.
Learning to play in a day: Faster deep reinforcement learning by optimality tightening
He, Frank S., Yang Liu, Alexander G. Schwing, and Peng, Jian · 2016
Closest in time.
State of the art control of Atari games using shallow reinforcement learning
Liang, Yitao, Machado, Marlos C, Talvitie, Erik, and Bowling, Michael · 2016
Closest in time.
Asynchronous methods for deep reinforcement learning
Mnih, Volodymyr, Badia, Adria Puigdomenech, Mirza, Mehdi, Graves, Alex, Lillicrap, Timothy P, Harley, Tim, Silver, David, and Kavukcuoglu, Koray · 2016
Closest in time.
Deep exploration via bootstrapped DQN
Osband, Ian, Blundell, Charles, Pritzel, Alexander, and Van Roy, Benjamin · 2016
Closest in time.
#exploration: A study of count-based exploration for deep reinforcement learning
Tang, Haoran, Rein Houthooft, Davis Foote, Adam Stooke, Xi Chen, Yan Duan, John Schulman, and Filip De Turck, Pieter Abbeel · 2016
Closest in time.