Fetching the paper…
Reading the bibliography…
Many deep reinforcement learning algorithms contain inductive biases that sculpt the agent's objective and its interface to the environment.
Dynamic programming
R. Bellman · 1957
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton · 1988
Earlier work this paper cites.
Adapting bias by gradient descent: An incremental version of delta-bar-delta
Sutton · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams · 1992
Earlier work this paper cites.
Feudal reinforcement learning
P. Dayan and G. E. Hinton · 1993
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
HQ-learning
M. Wiering and J. Schmidhuber · 1997
Earlier work this paper cites.
Intra-option learning about temporally abstract actions
Sutton., Precup, and Singh · 1998
Earlier work this paper cites.
Gradient-based optimization of hyperparameters
Y. Bengio · 2000
Earlier work this paper cites.
Policy gradient methods for rl with function approximation
Sutton, McAllester, Singh, and Mansour · 2000
Earlier work this paper cites.
Core knowledge
E. S. Spelke and K. D. Kinzler · 2007
Earlier work this paper cites.
Action-gap phenomenon in reinforcement learning
A.-m. Farahmand · 2011
Earlier work this paper cites.
Understanding the exploding gradient problem
R. Pascanu, T. Mikolov, and Y. Bengio · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
The Arcade Learning Environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Cited alongside, same era.
Increasing the action gap: New operators for reinforcement learning
M. G. Bellemare, G. Ostrovski, A. Guez, P. S. Thomas, and R. Munos · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Cited alongside, same era.
Deep reinforcement learning with double Q-learning
H. van Hasselt, A. Guez, and D. Silver · 2016
Later among the works it cites.
Dueling network architectures for deep reinforcement learning
Z. Wang, T. Schaul, M. Hessel, H. van Hasselt, M. Lanctot, and N. de Freitas · 2016
Later among the works it cites.
The option-critic architecture
P. Bacon, J. Harb, and D. Precup · 2017
Later among the works it cites.
Model-agnostic meta-learning for fast adaptation of deep networks
C. Finn, P. Abbeel, and S. Levine · 2017
Later among the works it cites.
Dynamic action repetition for deep reinforcement learning
Lakshminarayanan, Sharma, and Ravindran · 2017
Later among the works it cites.
Revisiting the Arcade Learning Environment: Evaluation Protocols and Open Problems for General Agents
M. C. Machado, M. G. Bellemare, E. Talvitie, J. Veness, M. Hausknecht, and M. Bowling · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Beattie, J. Z. Leibo, D. Teplyashin, T. Ward, M. Wainwright, H. Küttler, A. Lefrancq, S. Green, V. Valdés, A. Sadik, J. Schrittwieser, K. Anderson, S. York, M. Cant, A. Cain, A. Bolton, S. Gaffney, H. King, D. Hassabis, S. Legg, and S. Petersen · 2016
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
M. Jaderberg, V. Mnih, W. M. Czarnecki, T. Schaul, J. Z. Leibo, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
The Malmo platform for artificial intelligence experimentation
M. Johnson, K. Hofmann, T. Hutton, and D. Bignell · 2016
Cited alongside, same era.
Vizdoom: A doom-based ai research platform for visual rl
Kempka, Wydmuch, Runc, Toczek, and Jaśkowski · 2016
Cited alongside, same era.
Learning hand-eye coordination for robotic grasping with large-scale data collection
S. Levine, P. Pastor, A. Krizhevsky, and D. Quillen · 2016
Cited alongside, same era.
Asynchronous methods for deep rl
Mnih, Badia, Mirza, Graves, Lillicrap, Harley, Silver, and Kavukcuoglu · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Cited alongside, same era.
Later among the works it cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
Silver, Hubert, Schrittwieser, Antonoglou, Lai, Guez, Lanctot, Sifre, Kumaran, Graepel, Lillicrap, Simonyan, and Hassabis · 2017
Later among the works it cites.
Investigating human priors for playing video games
Dubey, Agrawal, Pathak, Griffiths, and Efros · 2018
Later among the works it cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, S. Legg, and K. Kavukcuoglu · 2018
Later among the works it cites.
Distributed prioritized experience replay
D. Horgan, J. Quan, D. Budden, G. Barth-Maron, M. Hessel, H. van Hasselt, and D. Silver · 2018
Later among the works it cites.
Innateness, alphazero, and artificial intelligence
Marcus · 2018
Later among the works it cites.
Control suite
Tassa, Doron, Muldal, Erez, Li, L. Casas, Budden, Abdolmaleki, Merel, Lefrancq, Lillicrap, and Riedmiller · 2018
Later among the works it cites.
Meta-gradient reinforcement learning
Xu, van Hasselt, and Silver · 2018
Later among the works it cites.