Fetching the paper…
Reading the bibliography…
Deep reinforcement learning methods attain super-human performance in a wide range of environments.
Multidimensional binary search trees used for associative searching
Bentley, Jon Louis · 1975
Earlier work this paper cites.
Using fast weights to deblur old memories
Hinton, Geoffrey E and Plaut, David C · 1987
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, Richard S · 1988
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
McCloskey, Michael and Cohen, Neal J · 1989
Earlier work this paper cites.
Learning from delayed rewards
Watkins, Christopher John Cornish Hellaby · 1989
Earlier work this paper cites.
Q-learning
Watkins, Christopher JCH and Dayan, Peter · 1992
Earlier work this paper cites.
Incremental multi-step q-learning
Peng, Jing and Williams, Ronald J · 1996
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Experiments with reinforcement learning in problems with continuous state and action spaces
Santamaría, Juan C, Sutton, Richard S, and Ram, Ashwin · 1997
Earlier work this paper cites.
Barycentric interpolators for continuous space and time reinforcement learning
Munos, Remi and Moore, Andrew W · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, Richard S and Barto, Andrew G · 1998
Earlier work this paper cites.
A robot that reinforcement-learns to identify and memorize important previous observations
Bakker, Bram, Zhumatiy, Viktor, Gruener, Gabriel, and Schmidhuber, Jürgen · 2003
Earlier work this paper cites.
Cbr for state value function approximation in reinforcement learning
Gabel, Thomas and Riedmiller, Martin · 2005
Earlier work this paper cites.
Hippocampal contributions to control: The third way
Lengyel, M. and Dayan, P · 2007
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, Tijmen and Hinton, Geoffrey · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Bellemare, Marc G, Naddaf, Yavar, Veness, Joel, and Bowling, Michael · 2013
Cited alongside, same era.
Auto-encoding variational bayes
Kingma, Diederik P and Welling, Max · 2013
Cited alongside, same era.
Stochastic backpropagation and approximate inference in deep generative models
Rezende, Danilo Jimenez, Mohamed, Shakir, and Wierstra, Daan · 2014
Cited alongside, same era.
Deep recurrent q-learning for partially observable mdps
Hausknecht, Matthew and Stone, Peter · 2015
Cited alongside, same era.
What learning systems do intelligent agents need? complementary learning systems theory updated
Kumaran, Dharshan, Hassabis, Demis, and McClelland, James L · 2016
Later among the works it cites.
Building machines that learn and think like people
Lake, Brenden M, Ullman, Tomer D, Tenenbaum, Joshua B, and Gershman, Samuel J · 2016
Later among the works it cites.
Key-value memory networks for directly reading documents
Miller, Alexander, Fisch, Adam, Dodge, Jesse, Karimi, Amir-Hossein, Bordes, Antoine, and Weston, Jason · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Mnih, Volodymyr, Badia, Adria Puigdomenech, Mirza, Mehdi, Graves, Alex, Lillicrap, Timothy P, Harley, Tim, Silver, David, and Kavukcuoglu, Koray · 2016
Later among the works it cites.
Safe and efficient off-policy reinforcement learning
Munos, Rémi, Stepleton, Tom, Harutyunyan, Anna, and Bellemare, Marc · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A, Veness, Joel, Bellemare, Marc G, Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K, Ostrovski, Georg, et al · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
Oh, Junhyuk, Guo, Xiaoxiao, Lee, Honglak, Lewis, Richard L, and Singh, Satinder · 2015
Cited alongside, same era.
End-to-end memory networks
Sukhbaatar, Sainbayar, Weston, Jason, Fergus, Rob, et al · 2015
Cited alongside, same era.
Using fast weights to attend to the recent past
Ba, Jimmy, Hinton, Geoffrey E, Mnih, Volodymyr, Leibo, Joel Z, and Ionescu, Catalin · 2016
Cited alongside, same era.
Blundell, Charles, Uria, Benigno, Pritzel, Alexander, Li, Yazhe, Ruderman, Avraham, Leibo, Joel Z, Rae, Jack, Wierstra, Daan, and Hassabis, Demis · 2016
Cited alongside, same era.
Rl 2 : Fast reinforcement learning via slow reinforcement learning
Duan, Yan, Schulman, John, Chen, Xi, Bartlett, Peter L, Sutskever, Ilya, and Abbeel, Pieter · 2016
Cited alongside, same era.
Hybrid computing using a neural network with dynamic external memory
Graves, Alex, Wayne, Greg, Reynolds, Malcolm, Harley, Tim, Danihelka, Ivo, Grabska-Barwińska, Agnieszka, Colmenarejo, Sergio Gómez, Grefenstette, Edward, Ramalho, Tiago, Agapiou, John, et al · 2016
Cited alongside, same era.
Later among the works it cites.
Control of memory, active perception, and action in minecraft
Oh, Junhyuk, Chockalingam, Valliappa, Lee, Honglak, et al · 2016
Later among the works it cites.
Deep exploration via bootstrapped dqn
Osband, Ian, Blundell, Charles, Pritzel, Alexander, and Van Roy, Benjamin · 2016
Later among the works it cites.
Rusu, Andrei A, Rabinowitz, Neil C, Desjardins, Guillaume, Soyer, Hubert, Kirkpatrick, James, Kavukcuoglu, Koray, Pascanu, Razvan, and Hadsell, Raia · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
Silver, David, Huang, Aja, Maddison, Chris J, Guez, Arthur, Sifre, Laurent, Van Den Driessche, George, Schrittwieser, Julian, Antonoglou, Ioannis, Panneershelvam, Veda, Lanctot, Marc, et al · 2016
Later among the works it cites.
Learning functions across many orders of magnitudes
van Hasselt, H., Guez, A., Hessel, M., and Silver, D · 2016
Later among the works it cites.
Deep reinforcement learning with double q-learning
Van Hasselt, Hado, Guez, Arthur, and Silver, David · 2016
Later among the works it cites.
Strategic attentive writer for learning macro-actions
Vezhnevets, Alexander, Mnih, Volodymyr, Osindero, Simon, Graves, Alex, Vinyals, Oriol, Agapiou, John, et al · 2016
Later among the works it cites.
Matching networks for one shot learning
Vinyals, Oriol, Blundell, Charles, Lillicrap, Tim, Wierstra, Daan, et al · 2016
Later among the works it cites.
Learning to reinforcement learn
Wang, Jane X, Kurth-Nelson, Zeb, Tirumala, Dhruva, Soyer, Hubert, Leibo, Joel Z, Munos, Remi, Blundell, Charles, Kumaran, Dharshan, and Botvinick, Matt · 2016
Later among the works it cites.
Pathnet: Evolution channels gradient descent in super neural networks
Fernando, Chrisantha, Banarse, Dylan, Blundell, Charles, Zwols, Yori, Ha, David, Rusu, Andrei A, Pritzel, Alexander, and Wierstra, Daan · 2017
Closest in time.