Fetching the paper…
Reading the bibliography…
Inspired by the phenomenon of catastrophic forgetting, we investigate the learning dynamics of neural networks as they train on single classification tasks.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J Cohen · 1989
Earlier work this paper cites.
Robust decision trees: removing outliers from databases
George H John · 1995
Earlier work this paper cites.
Flat minima
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Identifying mislabeled training data
Carla E Brodley and Mark A Friedl · 1999
Earlier work this paper cites.
The mnist database of handwritten digits
Y. LeCun, C. Cortes C., and C. Burges · 1999
Earlier work this paper cites.
Scaling learning algorithms towards AI
Yoshua Bengio and Yann LeCun · 2007
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Self-Paced Learning for Latent Variable Models
M Pawan Kumar, Benjamin Packer, and Daphne Koller · 2010
Earlier work this paper cites.
Learning the easy things first: Self-paced visual category discovery
Yong Jae Lee and Kristen Grauman · 2011
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
Training convolutional networks with noisy labels
Sainbayar Sukhbaatar, Joan Bruna, Manohar Paluri, Lubomir Bourdev, and Rob Fergus · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2014
Diederik Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2015
Earlier work this paper cites.
Stochastic Optimization with Importance Sampling for Regularized Loss Minimization
Peilin Zhao and Tong Zhang · 2015
Cited alongside, same era.
Entropy-SGD: Biasing Gradient Descent Into Wide Valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
Sergey Zagoruyko and Nikos Komodakis · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang · 2017
Later among the works it cites.
Optimization as a model for few-shot learning
Sachin Ravi and Hugo Larochelle · 2017
Later among the works it cites.
The Implicit Bias of Gradient Descent on Separable Data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2017
Later among the works it cites.
MentorNet: Learning data-driven curriculum for very deep neural networks on corrupted labels
Lu Jiang, Zhengyuan Zhou, Thomas Leung, Li-Jia Li, and Li Fei-Fei · 2018
Closest in time.
Not all samples are created equal: Deep learning with importance sampling
Angelos Katharopoulos and François Fleuret · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks
Madhu S. Advani and Andrew M. Saxe · 2017
Cited alongside, same era.
A closer look at memorization in deep networks
Devansh Arpit, Stanislaw Jastrzebski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, et al · 2017
Cited alongside, same era.
Active Bias: Training More Accurate Neural Networks by Emphasizing High Variance Samples
Haw-Shiuan Chang, Erik Learned-Miller, and Andrew McCallum · 2017
Cited alongside, same era.
Improved regularization of convolutional neural networks with cutout
Terrance DeVries and Graham W Taylor · 2017
Cited alongside, same era.
Yang Fan, Fei Tian, Tao Qin, and Jiang Bian · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Densely Connected Convolutional Networks
Gao Huang, Zhuang Liu, Laurens van der Maaten, and Kilian Q Weinberger · 2017
Cited alongside, same era.
Tae-Hoon Kim and Jonghyun Choi · 2018
Closest in time.
An alternative view: When does sgd escape local minima?
Robert Kleinberg, Yuanzhi Li, and Yang Yuan · 2018
Closest in time.
Measuring the intrinsic dimension of objective landscapes
Chunyuan Li, Heerad Farkhoor, Rosanne Liu, and Jason Yosinski · 2018
Closest in time.
Deep learning generalizes because the parameter-function map is biased towards simple functions
Guillermo Valle Perez, Chico Q. Camargo, and Ard A. Louis · 2018
Closest in time.
Online Structured Laplace Approximations For Overcoming Catastrophic Forgetting
Hippolyt Ritter, Aleksandar Botev, and David Barber · 2018
Closest in time.
On the learning dynamics of deep neural networks
R. Tachet, M. Pezeshki, S. Shabanian, A. Courville, and Y. Bengio · 2018
Closest in time.
Identifying Generalization Properties in Neural Networks
Huan Wang, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher · 2018
Closest in time.
Convergence of sgd in learning relu models with separable data
Tengyu Xu, Yi Zhou, Kaiyi Ji, and Yingbin Liang · 2018
Closest in time.