Fetching the paper…
Reading the bibliography…
Catastrophic forgetting affects the training of neural networks, limiting their ability to learn multiple tasks sequentially.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J. Cohen · 1989
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Using hindsight to anchor past knowledge in continual learning
Arslan Chaudhry, Albert Gordo, Puneet Kumar Dokania, Philip H. S. Torr, and David Lopez-Paz · 2002
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
Geoffrey E. Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2012
Earlier work this paper cites.
Understanding dropout
Pierre Baldi and Peter J Sadowski · 2013
Earlier work this paper cites.
An empirical investigation of catastrophic forgeting in gradient-based neural networks
Ian J. Goodfellow, Mehdi Mirza, Xia Da, Aaron C. Courville, and Yoshua Bengio · 2013
Earlier work this paper cites.
The stability-plasticity dilemma: investigating the continuum from catastrophic forgetting to age-limited learning effects
Martial Mermillod, Aurélia Bugaiska, and Patrick Bonin · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
Improving neural networks with dropout
Nitish Srivastava · 2013
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Ian J Goodfellow, Oriol Vinyals, and Andrew M Saxe · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Surprising properties of dropout in deep networks
David P. Helmbold and Philip M. Long · 2016
Earlier work this paper cites.
icarl: Incremental classifier and representation learning
Sylvestre-Alvise Rebuffi, Alexander I Kolesnikov, Georg Sperl, and Christoph H. Lampert · 2016
Earlier work this paper cites.
Andrei A Rusu, Neil C Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell · 2016
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Earlier work this paper cites.
Improving generalization performance by switching from adam to sgd
Nitish Shirish Keskar and Richard Socher · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James N Kirkpatrick, Razvan Pascanu, Neil C. Rabinowitz, Joel Veness, and et. al · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting by incremental moment matching
Sang-Woo Lee, Jin-Hwa Kim, Jaehyun Jun, Jung-Woo Ha, and Byoung-Tak Zhang · 2017
Earlier work this paper cites.
Gradient episodic memory for continual learning
David Lopez-Paz and Marc’Aurelio Ranzato · 2017
Earlier work this paper cites.
Packnet: Adding multiple tasks to a single network by iterative pruning
Arun Mallya and Svetlana Lazebnik · 2017
Earlier work this paper cites.
Variational continual learning
Cuong V Nguyen, Yingzhen Li, Thang D Bui, and Richard E Turner · 2017
Earlier work this paper cites.
Continual learning with deep generative replay
Hanul Shin, Jung Kwon Lee, Jaehong Kim, and Jiwon Kim · 2017
Earlier work this paper cites.
Don’t decay the learning rate, increase the batch size
Samuel L. Smith, Pieter-Jan Kindermans, and Quoc V. Le · 2017
Cited alongside, same era.
The marginal value of adaptive gradient methods in machine learning
Ashia C. Wilson, Rebecca Roelofs, Mitchell Stern, Nathan Srebro, and Benjamin Recht · 2017
Cited alongside, same era.
Continual learning through synaptic intelligence
Friedemann Zenke, Ben Poole, and Surya Ganguli · 2017
Cited alongside, same era.
Memory aware synapses: Learning what (not) to forget
Rahaf Aljundi, Francesca Babiloni, Mohamed Elhoseiny, Marcus Rohrbach, and Tinne Tuytelaars · 2018
Cited alongside, same era.
Closing the generalization gap of adaptive gradient methods in training deep neural networks
Jinghui Chen and Quanquan Gu · 2018
Cited alongside, same era.
Uncertainty-guided continual learning with bayesian neural networks
Sayna Ebrahimi, Mohamed Elhoseiny, Trevor Darrell, and Marcus Rohrbach · 2019
Later among the works it cites.
Orthogonal gradient descent for continual learning
Mehrdad Farajtabar, Navid Azizan, Alex Mott, and Ang Li · 2019
Later among the works it cites.
Emergent properties of the local geometry of neural loss landscapes
Stanislav Fort and Surya Ganguli · 2019
Later among the works it cites.
Demystifying dropout
Hongchang Gao, Jian Pei, and Heng Huang · 2019
Later among the works it cites.
The step decay schedule: A near optimal, geometrically decaying learning rate procedure
Rong Ge, Sham M. Kakade, Rahul Kidambi, and Praneeth Netrapalli · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards robust evaluations of continual learning
Sebastian Farquhar and Yarin Gal · 2018
Cited alongside, same era.
Dynamic few-shot visual learning without forgetting
Spyros Gidaris and Nikos Komodakis · 2018
Cited alongside, same era.
Overcoming catastrophic interference using conceptor-aided backpropagation
Xu He and Herbert Jaeger · 2018
Cited alongside, same era.
Re-evaluating continual learning scenarios: A categorization and case for strong baselines
Yen-Chang Hsu, Yen-Cheng Liu, and Zsolt Kira · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Measuring catastrophic forgetting in neural networks
Ronald Kemker, Marc McClure, Angelina Abitino, Tyler L Hayes, and Christopher Kanan · 2018
Cited alongside, same era.
Alleviating catastrophic forgetting using context-dependent gating and synaptic stabilization
Nicolas Y. Masse, Gregory D. Grant, and David J. Freedman · 2018
Cited alongside, same era.
Later among the works it cites.
Task agnostic continual learning via meta learning
Xu He, Jakub Sygnowski, Alexandre Galashov, Andrei A Rusu, Yee Whye Teh, and Razvan Pascanu · 2019
Later among the works it cites.
Reconciling meta-learning and continual learning with online mixtures of tasks
Ghassen Jerfel, Erin Grant, Thomas L. Griffiths, and Katherine A. Heller · 2019
Later among the works it cites.
Policy consolidation for continual reinforcement learning
Christos Kaplanis, Murray Shanahan, and Claudia Clopath · 2019
Later among the works it cites.
Attention-based structural-plasticity
Soheil Kolouri, Nicholas Ketz, Xinyun Zou, Jeffrey Krichmar, and Praveen Pilly · 2019
Later among the works it cites.
Continual learning: A comparative study on how to defy forgetting in classification tasks
Matthias Lange, Rahaf Aljundi, Marc Masana, Sarah Parisot, Xu Jia, Ale Leonardis, Gregory G. Slabaugh, and Tinne Tuytelaars · 2019
Later among the works it cites.
Learn to grow: A continual structure learning framework for overcoming catastrophic forgetting
Xilai Li, Yingbo Zhou, Tianfu Wu, Richard Socher, and Caiming Xiong · 2019
Later among the works it cites.
An exponential learning rate schedule for deep learning
Zhongyuan Li and Sanjeev Arora · 2019
Later among the works it cites.
Toward understanding catastrophic forgetting in continual learning
Cuong V Nguyen, Alessandro Achille, Michael Lam, Tal Hassner, Vijay Mahadevan, and Stefano Soatto · 2019
Later among the works it cites.
Continual unsupervised representation learning
Dushyant Rao, Francesco Visin, Andrei Rusu, Razvan Pascanu, Yee Whye Teh, and Raia Hadsell · 2019
Later among the works it cites.
Functional regularisation for continual learning using gaussian processes
Michalis K Titsias, Jonathan Schwarz, Alexander G de G Matthews, Razvan Pascanu, and Yee Whye Teh · 2019
Later among the works it cites.
Which algorithmic choices matter at which batch sizes? insights from a noisy quadratic model
Guodong Zhang, Lala Li, Zachary Nado, James Martens, Sushant Sachdeva, George Dahl, Chris Shallue, and Roger Grosse · 2019
Later among the works it cites.
Prototype reminding for continual learning
Mengmi Zhang, Tao Wang, Joo Hwee Lim, and Jiashi Feng · 2019
Later among the works it cites.
Shawn Beaulieu, Lapo Frati, Thomas Miconi, Joel Lehman, Kenneth O Stanley, Jeff Clune, and Nick Cheney · 2020
Closest in time.
The early phase of neural network training
Jonathan Frankle, David J. Schwab, and Ari S. Morcos · 2020
Closest in time.
The break-even point on the optimization trajectories of deep neural networks
Stanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit, Jacek Tabor, Kyunghyun Cho, and Krzysztof Geras · 2020
Closest in time.
The large learning rate phase of deep learning: the catapult mechanism
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2020
Closest in time.
Dropout as an implicit gating mechanism for continual learning
Seyed Iman Mirzadeh, Mehrdad Farajtabar, and Hassan Ghasemzadeh · 2020
Closest in time.
The implicit and explicit regularization effects of dropout
Colin Wei, Sham M. Kakade, and Tengyu Ma · 2020
Closest in time.
Zeke Xie, Issei Sato, and Masashi Sugiyama · 2020
Closest in time.