J. Schmidhuber, “Evolutionary principles in self-referential learning,” Diploma thesis, Tech. Univ. Munich , 1987
1987
Earlier work this paper cites.
M. McCloskey and N. J. Cohen, “Catastrophic interference in connectionist networks: The sequential learning problem,” in Psychology of learning and motivation , 1989
1989
Earlier work this paper cites.
R. Ratcliff, “Connectionist models of recognition memory: constraints imposed by learning and forgetting functions.” Psychological review , 1990
1990
Earlier work this paper cites.
A. W. Moore and C. G. Atkeson, “Prioritized sweeping: Reinforcement learning with less data and less time,” Machine learning , 1993
1993
Earlier work this paper cites.
S. Thrun, “A lifelong learning perspective for mobile robot control,” in Intelligent Robots and Systems , 1995
1995
Earlier work this paper cites.
S. Thrun, “Is learning the n-th thing any easier than learning the first?” in Proc. Adv. Neural Information Processing Systems , 1996
1996
Earlier work this paper cites.
R. M. French, “Catastrophic forgetting in connectionist networks,” Trends in cognitive sciences , 1999
1999
Earlier work this paper cites.
D. L. Silver and R. E. Mercer, “The task rehearsal method of life-long learning: Overcoming impoverished data,” in Conference of the Canadian Society for Computational Studies of Intelligence , 2002
2002
Earlier work this paper cites.
C. Buciluǎ, R. Caruana, and A. Niculescu-Mizil, “Model compres-sion,” in International Conference on Knowledge Discovery and Data Mining , 2006
2006
Earlier work this paper cites.
M.-E. Nilsback and A. Zisserman, “Automated flower classification over a large number of classes,” in Indian Conference on Computer Vision, Graphics & Image Processing , 2008
2008
Earlier work this paper cites.
A. Torralba, R. Fergus, and W. T. Freeman, “80 million tiny images: A large data set for nonparametric object and scene recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2008
2008
Earlier work this paper cites.
M. Welling, “Herding dynamical weights to learn,” in Proc. International Conference on Machine Learning , 2009
2009
Earlier work this paper cites.
A. Krizhevsky, “Learning multiple layers of features from tiny images,” Citeseer, Tech. Rep., 2009
2009
Earlier work this paper cites.
A. Quattoni and A. Torralba, “Recognizing indoor scenes,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition , 2009
2009
Earlier work this paper cites.
Y. Bengio, J. Louradour, R. Collobert, and J. Weston, “Curriculum learning,” in Proc. International Conference on Machine Learning , 2009
2009
Earlier work this paper cites.
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie, “The caltech-ucsd birds-200-2011 dataset,” California Institute of Technology, Tech. Rep. CNS-TR-2011-001, 2011
2011
Earlier work this paper cites.
B. Yao, X. Jiang, A. Khosla, A. L. Lin, L. Guibas, and L. Fei-Fei, “Human action recognition by learning bases of action attributes and parts,” in Proc. IEEE International Conference on Computer Vision , 2011
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Proc. Adv. Neural Information Processing Systems , 2012
2012
Earlier work this paper cites.
M. Mermillod, A. Bugaiska, and P. Bonin, “The stability-plasticity dilemma: Investigating the continuum from catastrophic forgetting to age-limited learning effects,” Frontiers in psychology , 2013
2013
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller, “Playing atari with deep reinforcement learning,” in Proc. Adv. Neural Information Processing Systems Workshop on Deep Learning , 2013
2013
Earlier work this paper cites.
R. Coop and I. Arel, “Mitigation of catastrophic forgetting in recurrent neural networks using a fixed expansion layer,” in International Joint Conference on Neural Networks , 2013
2013
Earlier work this paper cites.
J. Krause, M. Stark, J. Deng, and L. Fei-Fei, “3d object representations for fine-grained categorization,” in Proc. IEEE International Conference on Computer Vision Workshops , 2013
2013
Earlier work this paper cites.
S. Maji, J. Kannala, E. Rahtu, M. Blaschko, and A. Vedaldi, “Fine-grained visual classification of aircraft,” eprint arXiv:1306.5151, Tech. Rep., 2013
Original
2013
Earlier work this paper cites.
I. J. Goodfellow, M. Mirza, D. Xiao, A. Courville, and Y. Bengio, “An empirical investigation of catastrophic forgetting in gradient-based neural networks,” in Proc. International Conference on Learning Representations , 2014
2014
Earlier work this paper cites.
G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” in Proc. Adv. Neural Information Processing Systems Workshop on Deep Learning , 2014
2014
Earlier work this paper cites.
T. Xiao, J. Zhang, K. Yang, Y. Peng, and Z. Zhang, “Error-driven incremental learning in deep convolutional neural network for large-scale image classification,” in ACM International Conference on Multimedia , 2014
2014
Earlier work this paper cites.
R. Girshick, “Fast r-cnn,” in IEEE International Conference on Computer Vision , 2015
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” in Proc. Adv. Neural Information Processing Systems , 2015
2015
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al. , “Imagenet large scale visual recognition challenge,” International Journal of Computer Vision , 2015
2015
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in Proc. International Conference on Learning Representations , 2015
2015
Earlier work this paper cites.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in IEEE Conference on Computer Vision and Pattern Recognition , 2015
2015
Earlier work this paper cites.
A. A. Rusu, N. C. Rabinowitz, G. Desjardins, H. Soyer, J. Kirkpatrick, K. Kavukcuoglu, R. Pascanu, and R. Hadsell, “Progressive neural networks,” arXiv , 2016
2016
Earlier work this paper cites.
H. Jung, J. Ju, M. Jung, and J. Kim, “Less-forgetting learning in deep neural networks,” arXiv , 2016
2016
Earlier work this paper cites.
T. Chen, I. Goodfellow, and J. Shlens, “Net2net: Accelerating learning via knowledge transfer,” in Proc. International Conference on Learning Representations , 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition , 2016
2016
Earlier work this paper cites.
S.-A. Rebuffi, A. Kolesnikov, G. Sperl, and C. H. Lampert, “icarl: Incremental classifier and representation learning,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition , 2017
2017
Earlier work this paper cites.
J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska et al. , “Overcoming catastrophic forgetting in neural networks,” National Academy of Sciences , 2017
2017
Earlier work this paper cites.
R. Aljundi, P. Chakravarty, and T. Tuytelaars, “Expert gate: Lifelong learning with a network of experts,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition , 2017
2017
Earlier work this paper cites.
D. Lopez-Paz and M. Ranzato, “Gradient episodic memory for continual learning,” in Proc. Adv. Neural Information Processing Systems , 2017
2017
Earlier work this paper cites.
Z. Li and D. Hoiem, “Learning without forgetting,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2017
2017
Earlier work this paper cites.
H. Shin, J. K. Lee, J. Kim, and J. Kim, “Continual learning with deep generative replay,” in Proc. Adv. Neural Information Processing Systems , 2017
2017
Earlier work this paper cites.
F. Zenke, B. Poole, and S. Ganguli, “Continual learning through synaptic intelligence,” in Proc. International Conference on Machine Learning , 2017
2017
Earlier work this paper cites.
S.-W. Lee, J.-H. Kim, J. Jun, J.-W. Ha, and B.-T. Zhang, “Overcoming catastrophic forgetting by incremental moment matching,” in Proc. Adv. Neural Information Processing Systems , 2017
2017
Earlier work this paper cites.
A. Rannen, R. Aljundi, M. B. Blaschko, and T. Tuytelaars, “Encoder based lifelong learning,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition , 2017
2017
Earlier work this paper cites.