Fetching the paper…
Reading the bibliography…
Experience replay lets online reinforcement learning agents remember and reuse experiences from the past.
Curious model-building control systems
Schmidhuber, Jürgen · 1991
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, Long-Ji · 1992
Earlier work this paper cites.
Q-learning
Watkins, Christopher JCH and Dayan, Peter · 1992
Earlier work this paper cites.
Prioritized sweeping: Reinforcement learning with less data and less time
Moore, Andrew W and Atkeson, Christopher G · 1993
Earlier work this paper cites.
Zebras and the Anna Karenina principle
Diamond, Jared · 1994
Earlier work this paper cites.
Rprop-description and implementation details
Riedmiller, Martin · 1994
Earlier work this paper cites.
New methods for competitive coevolution
Rosin, Christopher D and Belew, Richard K · 1997
Earlier work this paper cites.
Generalized prioritized sweeping
Andre, David, Friedman, Nir, and Parr, Ronald · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Lecun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
The MNIST database of handwritten digits, 1998
LeCun, Yann, Cortes, Corinna, and Burges, Christopher JC · 1998
Earlier work this paper cites.
Reverse replay of behavioural sequences in hippocampal place cells during the awake state
Foster, David J and Wilson, Matthew A · 2006
Earlier work this paper cites.
To recognize shapes, first learn to generate images
Hinton, Geoffrey E · 2007
Earlier work this paper cites.
A discriminatively trained, multiscale, deformable part model
Felzenszwalb, Pedro, McAllester, David, and Ramanan, Deva · 2008
Earlier work this paper cites.
Rewarded outcomes enhance reactivation of experience in the hippocampus
Singer, Annabelle C and Frank, Loren M · 2009
Cited alongside, same era.
Double Q-learning
van Hasselt, Hado · 2010
Cited alongside, same era.
Torch7: A matlab-like environment for machine learning
Collobert, Ronan, Kavukcuoglu, Koray, and Farabet, Clément · 2011
Cited alongside, same era.
Online discovery of feature dependencies
Geramifard, Alborz, Doshi, Finale, Redding, Joshua, Roy, Nicholas, and How, Jonathan · 2011
Cited alongside, same era.
Incremental basis construction from temporal difference error
Sun, Yi, Ring, Mark, Schmidhuber, Jürgen, and Gomez, Faustino J · 2011
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Bellemare, Marc G, Naddaf, Yavar, Veness, Joel, and Bowling, Michael · 2012
Weighted importance sampling for off-policy learning with linear function approximation
Mahmood, A Rupam, van Hasselt, Hado P, and Sutton, Richard S · 2014
Later among the works it cites.
Dopaminergic neurons promote hippocampal reactivation and spatial memory persistence
McNamara, Colin G, Tejero-Cantero, Álvaro, Trouche, Stéphanie, Campo-Urriza, Natalia, and Dupret, David · 2014
Later among the works it cites.
Surprise and curiosity for big data robotics
White, Adam, Modayil, Joseph, and Sutton, Richard S · 2014
Later among the works it cites.
Memory trace replay: the shaping of memory consolidation by neuromodulation
Atherton, Laura A, Dupret, David, and Mellor, Jack R · 2015
Closest in time.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A, Veness, Joel, Bellemare, Marc G, Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K, Ostrovski, Georg, Petersen, Stig, Beattie, Charles, Sadik, Amir, Antonoglou, Ioannis, King, Helen, Kumaran, Dharshan, Wierstra, Daan, Legg, Shane, and Hassabis, Demis · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A review on ensembles for the class imbalance problem: bagging-, boosting-, and hybrid-based approaches
Galar, Mikel, Fernandez, Alberto, Barrenechea, Edurne, Bustince, Humberto, and Herrera, Francisco · 2012
Cited alongside, same era.
Playing atari with deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Graves, Alex, Antonoglou, Ioannis, Wierstra, Daan, and Riedmiller, Martin · 2013
Cited alongside, same era.
No more pesky learning rates
Schaul, Tom, Zhang, Sixin, and Lecun, Yann · 2013
Cited alongside, same era.
Planning by prioritized sweeping with small backups
van Seijen, Harm and Sutton, Richard · 2013
Cited alongside, same era.
Deep Learning for Real-Time Atari Game Play Using Offline Monte-Carlo Tree Search Planning
Guo, Xiaoxiao, Singh, Satinder, Lee, Honglak, Lewis, Richard L, and Wang, Xiaoshi · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, Diederik P. and Ba, Jimmy · 2014
Cited alongside, same era.
Massively parallel methods for deep reinforcement learning
Nair, Arun, Srinivasan, Praveen, Blackwell, Sam, Alcicek, Cagdas, Fearon, Rory, Maria, Alessandro De, Panneershelvam, Vedavyas, Suleyman, Mustafa, Beattie, Charles, Petersen, Stig, Legg, Shane, Mnih, Volodymyr, Kavukcuoglu, Koray, and Silver, David · 2015
Closest in time.
Language understanding for text-based games using deep reinforcement learning
Narasimhan, Karthik, Kulkarni, Tejas, and Barzilay, Regina · 2015
Closest in time.
Hippocampal place cells construct reward related sequences through unexplored space
Ólafsdóttir, H Freyja, Barry, Caswell, Saleem, Aman B, Hassabis, Demis, and Spiers, Hugo J · 2015
Closest in time.
Incentivizing exploration in reinforcement learning with deep predictive models
Stadie, Bradly C, Levine, Sergey, and Abbeel, Pieter · 2015
Closest in time.
Dueling network architectures for deep reinforcement learning
Wang, Z., de Freitas, N., and Lanctot, M · 2015
Closest in time.
Increasing the action gap: New operators for reinforcement learning
Bellemare, Marc G., Ostrovski, Georg, Guez, Arthur, Thomas, Philip S., and Munos, Rémi · 2016
Closest in time.
Deep Reinforcement Learning with Double Q-learning
van Hasselt, Hado, Guez, Arthur, and Silver, David · 2016
Closest in time.