Fetching the paper…
Reading the bibliography…
This work identifies a common flaw of deep reinforcement learning (RL) algorithms: a tendency to rely on early interactions and ignore useful evidence encountered later.
Forming impressions of personality
Asch, S. E · 1961
Earlier work this paper cites.
The effects of the elimination of rehearsal on primacy and recency
Marshall, P. H. and Werder, P. R · 1972
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Critical period effects in second language learning: The influence of maturational state on the acquisition of english as a second language
Johnson, J. S. and Newport, E. L · 1989
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L. J · 1992
Earlier work this paper cites.
Q-learning with hidden-unit restarting
Anderson, C. W. et al · 1993
Earlier work this paper cites.
Catastrophic forgetting in neural networks: the role of rehearsal mechanisms
Robins, A · 1993
Earlier work this paper cites.
An analysis of catastrophic interference
Sharkey, N. E. and Sharkey, A. J · 1995
Earlier work this paper cites.
Python tutorial , volume 620
Van Rossum, G. and Drake Jr, F. L · 1995
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
French, R. M · 1999
Earlier work this paper cites.
Spontaneous evolution of linguistic structure-an iterated learning model of the emergence of regularity and irregularity
Kirby, S · 2001
Earlier work this paper cites.
A guide to NumPy , volume 1
Oliphant, T. E · 2006
Earlier work this paper cites.
Matplotlib: A 2d graphics environment
Hunter, J. D · 2007
Earlier work this paper cites.
Python for scientific computing
Oliphant, T. E · 2007
Earlier work this paper cites.
Regularized policy iteration
Farahmand, A., Ghavamzadeh, M., Mannor, S., and Szepesvári, C · 2008
Earlier work this paper cites.
Why does unsupervised pre-training help deep learning?
Erhan, D., Courville, A., Bengio, Y., and Vincent, P · 2010
Earlier work this paper cites.
The numpy array: a structure for efficient numerical computation
Van Der Walt, S., Colbert, S. C., and Varoquaux, G · 2011
Earlier work this paper cites.
Python for data analysis: Data wrangling with Pandas, NumPy, and IPython
McKinney, W · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
The role of first impression in operant learning
Shteingart, H., Neiman, T., and Loewenstein, Y · 2013
Earlier work this paper cites.
SciPy: Open source scientific tools for Python
Jones, E., Oliphant, T., and Peterson, P · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
A deeper look at planning as learning from replay
Vanseijen, H. and Sutton, R · 2015
Cited alongside, same era.
Jupyter Notebooks-a publishing format for reproducible computational workflows. , volume 2016
Kluyver, T., Ragan-Kelley, B., Pérez, F., Granger, B. E., Bussonnier, M., Frederic, J., Kelley, K., Hamrick, J. B., Grout, J., Corlay, S., et al · 2016
Cited alongside, same era.
Prioritized experience replay
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2016
Revisiting fundamentals of experience replay
Fedus, W., Ramachandran, P., Agarwal, R., Bengio, Y., Larochelle, H., Rowland, M., and Dabney, W · 2020
Later among the works it cites.
Implicit under-parameterization inhibits data-efficient deep reinforcement learning
Kumar, A., Agarwal, R., Ghosh, D., and Levine, S · 2020
Later among the works it cites.
Data-efficient reinforcement learning with self-predictive representations
Schwarzer, M., Anand, A., Goel, R., Hjelm, R. D., Courville, A., and Bachman, P · 2020
Later among the works it cites.
dm_control: Software and tasks for continuous control, 2020
Tassa, Y., Tunyasuvunakool, S., Muldal, A., Doron, Y., Liu, S., Bohez, S., Merel, J., Erez, T., Lillicrap, T., and Heess, N · 2020
Later among the works it cites.
Striving for simplicity and performance in off-policy drl: Output normalization and non-uniform sampling
Wang, C., Wu, Y., Vuong, Q., and Ross, K · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Distral: Robust multitask reinforcement learning
Teh, Y. W., Bapst, V., Czarnecki, W. M., Quan, J., Kirkpatrick, J., Hadsell, R., Heess, N., and Pascanu, R · 2017
Cited alongside, same era.
Critical learning periods in deep networks
Achille, A., Rovere, M., and Soatto, S · 2018
Cited alongside, same era.
What doubling tricks can and can’t do for multi-armed bandits
Besson, L. and Kaufmann, E · 2018
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Cited alongside, same era.
Soft actor-critic algorithms and applications
Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., et al · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D · 2018
Cited alongside, same era.
Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A., and Bellemare, M. G · 2021
Later among the works it cites.
The impact of reinitialization on generalization in convolutional neural networks
Alabdulmohsin, I., Maennel, H., and Keysers, D · 2021
Later among the works it cites.
Is high variance unavoidable in rl? a case study in continuous control
Bjorck, J., Gomes, C. P., and Weinberger, K. Q · 2021
Later among the works it cites.
The value-improvement path: Towards better representations for reinforcement learning
Dabney, W., Barreto, A., Rowland, M., Dadashi, R., Quan, J., Bellemare, M. G., and Silver, D · 2021
Later among the works it cites.
Continual backprop: Stochastic gradient descent with persistent randomness
Dohare, S., Mahmood, A. R., and Sutton, R. S · 2021
Later among the works it cites.
Transient non-stationarity and generalisation in deep reinforcement learning
Igl, M., Farquhar, G., Luketina, J., Boehmer, W., and Whiteson, S · 2021
Later among the works it cites.
A survey of generalisation in deep reinforcement learning
Kirk, R., Zhang, A., Grefenstette, E., and Rocktäschel, T · 2021
Later among the works it cites.
JAXRL: Implementations of Reinforcement Learning algorithms in JAX, 10 2021
Kostrikov, I · 2021
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Kostrikov, I., Yarats, D., and Fergus, R · 2021
Later among the works it cites.
On the convergence of stochastic extragradient for bilinear games with restarted iteration averaging
Li, C. J., Yu, Y., Loizou, N., Gidel, G., Ma, Y., Roux, N. L., and Jordan, M. I · 2021
Later among the works it cites.
Regularization matters in policy optimization - an empirical study on continuous control
Liu, Z., Li, X., Kang, B., and Darrell, T · 2021
Later among the works it cites.
Knowledge evolution in neural networks
Taha, A., Shrivastava, A., and Davis, L. S · 2021
Later among the works it cites.
Forgetting enhances episodic control with structured memories
Yalnizyan-Carson, A. and Richards, B. A · 2021
Later among the works it cites.
Improving sample efficiency in model-free reinforcement learning from images
Yarats, D., Zhang, A., Kostrikov, I., Amos, B., Pineau, J., and Fergus, R · 2021
Later among the works it cites.
Fortuitous forgetting in connectionist networks
Zhou, H., Vani, A., Larochelle, H., and Courville, A · 2021
Later among the works it cites.
Is high variance unavoidable in RL? a case study in continuous control
Bjorck, J., Gomes, C. P., and Weinberger, K. Q · 2022
Closest in time.
Understanding and preventing capacity loss in reinforcement learning
Lyle, C., Rowland, M., and Dabney, W · 2022
Closest in time.