Fetching the paper…
Reading the bibliography…
Modern deep-learning systems are specialized to problem settings in which training occurs once and then never again, as opposed to continual-learning settings in which training occurs continually.
A logical calculus of the ideas immanent in nervous activity
McCulloch, W. S. & Pitts, W · 1943
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E. & Williams, R. J · 1986
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
McCloskey, M. & Cohen, N. J · 1989
Earlier work this paper cites.
Optimal brain damage
LeCun, Y., Denker, J. & Solla, S · 1989
Earlier work this paper cites.
Using additive noise in back-propagation training
Holmstrom, L., Koistinen, P. et al · 1992
Earlier work this paper cites.
Learning in Embedded Systems (MIT Press, 1993)
Kaelbling, L. P · 1993
Earlier work this paper cites.
Online learning with random representations
Sutton, R. S. & Whitehead, S. D · 1993
Earlier work this paper cites.
Evaluating pruning methods
Thimm, G. & Fiesler, E · 1995
Earlier work this paper cites.
Multitask learning
Caruana, R · 1997
Earlier work this paper cites.
Child: A first step towards continual learning
Ring, M. B · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Lecun, Y., Bottou, L., Bengio, Y. & Haffner, P · 1998
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
French, R. M · 1999
Earlier work this paper cites.
Age of acquisition effects in adult lexical processing reflect loss of plasticity in maturing systems: insights from connectionist networks
Ellis, A. W. & Lambon Ralph, M. A · 2000
Earlier work this paper cites.
Age of acquisition effects in word reading and other tasks
Zevin, J. D. & Seidenberg, M. S · 2002
Earlier work this paper cites.
The influence of age of acquisition in word reading and other tasks: A never ending story?
Bonin, P., Barry, C., Méot, A. & Chalard, M · 2004
Earlier work this paper cites.
The effective rank: A measure of effective dimensionality
Roy, O. & Vetterli, M · 2007
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Deng, J. et al · 2009
Earlier work this paper cites.
Mnist handwritten digit database
LeCun, Y., Cortes, C. & Burges, C · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. & Bengio, Y · 2010
Earlier work this paper cites.
Rectified linear units improve Restricted Boltzmann Machines
Nair, V. & Hinton, G. E · 2010
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Hinton, G. et al · 2012
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I. & Hinton, G. E · 2012
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
Hinton, G. E., Srivastava, N., Krizhevsky, A., Sutskever, I. & Salakhutdinov, R. R · 2012
Earlier work this paper cites.
Neural Networks: Tricks of the Trade (Springer, 2012)
Montavon, G., Orr, G. & Müller, K.-R · 2012
Earlier work this paper cites.
Online incremental feature learning with denoising autoencoders
Zhou, G., Sohn, K. & Lee, H · 2012
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Pascanu, R., Mikolov, T. & Bengio, Y · 2013
Earlier work this paper cites.
Representation search through generate and test
Mahmood, A. R. & Sutton, R. S · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G. & Hinton, G · 2013
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Graves, A., Mohamed, A.-r. & Hinton, G · 2013
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
Maas, A. L., Hannun, A. Y. & Ng, A. Y · 2013
Earlier work this paper cites.
An empirical investigation of catastrophic forgeting in gradient-based neural networks
Goodfellow, I., Mirza, M., Xiao, D. & Aaron Courville, Y. B · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V. et al · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. & Ba, J · 2015
Earlier work this paper cites.
ImageNet large scale visual recognition challenge
Russakovsky, O. et al · 2015
Cited alongside, same era.
Measuring saturation in neural networks
Rakitianskaia, A. & Engelbrecht, A · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. & Szegedy, C · 2015
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S. & Sun, J · 2015
Cited alongside, same era.
Adding gradient noise improves learning for very deep networks
Neelakantan, A. et al · 2015
Cited alongside, same era.
Deep Learning (MIT Press, 2016)
Goodfellow, I., Bengio, Y. & Courville, A · 2016
Continual learning via neural pruning
Golkar, S., Kagan, M. & Cho, K · 2019
Later among the works it cites.
Learning to learn without forgetting by maximizing transfer and minimizing interference
Riemer, M. et al · 2019
Later among the works it cites.
Random path selection for continual learning
Rajasegaran, J., Hayat, M., Khan, S. H., Khan, F. S. & Shao, L · 2019
Later among the works it cites.
Meta-learning representations for continual learning
Javed, K. & White, M · 2019
Later among the works it cites.
Continual lifelong learning with neural networks: A review
Parisi, G. I., Kemker, R., Part, J. L., Kanan, C. & Wermter, S · 2019
Later among the works it cites.
Dying ReLU and initialization: Theory and numerical examples
Lu, L., Shin, Y., Su, Y. & Karniadakis, G. E · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Nonlinear Programming (Athena Scientific, 2016)
Bertsekas, D · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P. et al · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V. et al · 2016
Cited alongside, same era.
Network trimming: A data-driven neuron pruning approach towards efficient deep architectures
Hu, H., Peng, R., Tai, Y.-W. & Tang, C.-K · 2016
Cited alongside, same era.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
Han, S., Huizi, M. & Dally, W. J · 2016
Cited alongside, same era.
Rusu, A. A. et al · 2016
Cited alongside, same era.
Later among the works it cites.
Online normalization for training neural networks
Chiley, V. et al · 2019
Later among the works it cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. & Carbin, M · 2019
Later among the works it cites.
Learning to adapt in dynamic, real-world environments through meta-reinforcement learning
Nagabandi, A. et al · 2019
Later among the works it cites.
Language models are few-shot learners
Brown, T. et al · 2020
Later among the works it cites.
On warm-starting neural network training
Ash, J. & Adams, R. P · 2020
Later among the works it cites.
Trainability of relu networks and data-dependent initialization
Shin, Y. & Karniadakis, G. E · 2020
Later among the works it cites.
Implicit regularization in deep learning may not be explainable by norms
Razin, N. & Cohen, N · 2020
Later among the works it cites.
Dynamic sparse training: Find efficient sparse network from scratch with trainable masked layers
Liu, J., Xu, Z., Shi, R., Cheung, R. C. C. & So, H. K · 2020
Later among the works it cites.
Zero-shot text-to-image generation
Ramesh, A. et al · 2021
Later among the works it cites.
Efficient continual learning with modular networks and task-driven priors
Veniat, T., Denoyer, L. & Ranzato, M · 2021
Later among the works it cites.
A study on the plasticity of neural networks
Berariu, T. et al · 2021
Later among the works it cites.
On the origin of implicit regularization in stochastic gradient descent
Smith, S. L., Dherin, B., Barrett, D. & De, S · 2021
Later among the works it cites.
Chasing sparsity in vision transformers: An end-to-end exploration
Chen, T. et al · 2021
Later among the works it cites.
Transient non-stationarity and generalisation in deep reinforcement learning
Igl, M., Farquhar, G., Luketina, J., Boehmer, W. & Whiteson, S · 2021
Later among the works it cites.
Implicit under-parameterization inhibits data-efficient deep reinforcement learning
Kumar, A., Agarwal, R., Ghosh, D. & Levine, S · 2021
Later among the works it cites.
What matters for on-policy deep actor-critic methods? a large-scale study
Andrychowicz, M. et al · 2021
Later among the works it cites.
The primacy bias in deep reinforcement learning
Nikishin, E., Schwarzer, M., D’Oro, P., Bacon, P.-L. & Courville, A · 2022
Later among the works it cites.
Understanding and preventing capacity loss in reinforcement learning
Lyle, C., Rowland, M. & Dabney, W · 2022
Later among the works it cites.
Dynamic sparse training for deep reinforcement learning
Sokar, G., Mocanu, E., Mocanu, D. C., Pechenizkiy, M. & Stone, P · 2022
Later among the works it cites.
Understanding plasticity in neural networks
Lyle, C. et al · 2023
Closest in time.
Loss of plasticity in continual deep reinforcement learning
Abbas, Z., Zhao, R., Modayil, J., White, A. & Machado, M. C · 2023
Closest in time.
Continual learning as computationally constrained reinforcement learning
Kumar, S. et al · 2023
Closest in time.
Deep reinforcement learning with plasticity injection
Nikishin, E. et al · 2023
Closest in time.
Sample-efficient reinforcement learning by breaking the replay ratio barrier
D’Oro, P. et al · 2023
Closest in time.
The dormant neuron phenomenon in deep reinforcement learning
Sokar, G., Agarwal, R., Castro, P. S. & Evci, U · 2023
Closest in time.
Bigger, better, faster: Human-level Atari with human-level efficiency
Schwarzer, M. et al · 2023
Closest in time.