Fetching the paper…
Reading the bibliography…
Recent results suggest that reinitializing a subset of the parameters of a neural network during training can improve generalization, particularly for small training sets.
MNIST-C: A Robustness Benchmark for Computer Vision
Mu, N.; and Gilmer, J. 2019 · 1906
Earlier work this paper cites.
iCassava 2019 Fine-Grained Visual Categorization Challenge
Mwebaze, E.; Gebru, T.; Frome, A.; Nsumba, S.; and Tusubira, J. 2019 · 1908
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N.; Hinton, G.; Krizhevsky, A.; Sutskever, I.; and Salakhutdinov, R. 2014 · 1958
Earlier work this paper cites.
Training with noise is equivalent to Tikhonov regularization
Bishop, C. M. 1995 · 1995
Earlier work this paper cites.
Boosting the Margin: A New Explanation for the Effectiveness of Voting Methods
Schapire, R. E.; Freund, Y.; Bartlett, P. L.; and Lee, W. S. 1997 · 1997
Earlier work this paper cites.
The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network
Bartlett, P. L. 1998 · 1998
Earlier work this paper cites.
Learning Generative Visual Models from Few Training Examples: An Incremental Bayesian Approach Tested on 101 Object Categories
Fei-Fei, L.; Fergus, R.; and Perona, P. 2004 · 2004
Earlier work this paper cites.
Statistical comparisons of classifiers over multiple data sets
Demšar, J. 2006 · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. 2009 · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X.; and Bengio, Y. 2010 · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Nair, V.; and Hinton, G. E. 2010 · 2010
Earlier work this paper cites.
Caltech-UCSD Birds 200
Welinder, P.; Branson, S.; Mita, T.; Wah, C.; Schroff, F.; Belongie, S.; and Perona, P. 2010 · 2010
Earlier work this paper cites.
Novel Dataset for Fine-Grained Image Categorization
Khosla, A.; Jayadevaprakash, N.; Yao, B.; and Fei-Fei, L. 2011 · 2011
Earlier work this paper cites.
Scikit-learn: Machine Learning in Python
Pedregosa, F.; Varoquaux, G.; Gramfort, A.; Michel, V.; Thirion, B.; et al. 2011 · 2011
Earlier work this paper cites.
A Statistical-topological Feature Combination for Recognition of Handwritten Numerals
Das, N.; Reddy, J. M.; Sarkar, R.; Basu, S.; Kundu, M.; Nasipuri, M.; and Basu, D. K. 2012 · 2012
Earlier work this paper cites.
Cats and Dogs
Parkhi, O. M.; Vedaldi, A.; Zisserman, A.; and Jawahar, C. V. 2012 · 2012
Earlier work this paper cites.
3D Object Representations for Fine-Grained Categorization
Krause, J.; Stark, M.; Deng, J.; and Fei-Fei, L. 2013 · 2013
Cited alongside, same era.
How transferable are features in deep neural networks?
Yosinski, J.; Clune, J.; Bengio, Y.; and Lipson, H. 2014 · 2014
Cited alongside, same era.
TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems
Abadi, M.; Agarwal, A.; Barham, P.; Brevdo, E.; Chen, Z.; Citro, C.; Corrado, G. S.; Davis, A.; Dean, J.; Devin, M.; et al. 2015 · 2015
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2015 · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S.; and Szegedy, C. 2015 · 2015
Cited alongside, same era.
Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
Xiao, H.; Rasul, K.; and Vollgraf, R. 2017 · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Zhang, C.; Bengio, S.; Hardt, M.; Recht, B.; and Vinyals, O. 2017 · 2017
Later among the works it cites.
Stronger generalization bounds for deep nets via a compression approach
Arora, S.; Ge, R.; Neyshabur, B.; and Zhang, Y. 2018 · 2018
Later among the works it cites.
DNN or k-NN: That is the Generalize vs. Memorize Question
Cohen, G.; Sapiro, G.; and Giryes, R. 2018 · 2018
Later among the works it cites.
The implicit bias of gradient descent on separable data
Soudry, D.; Hoffer, E.; Nacson, M. S.; Gunasekar, S.; and Srebro, N. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Norm-based capacity control in neural networks
Neyshabur, B.; Tomioka, R.; and Srebro, N. 2015 · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Simonyan, K.; and Zisserman, A. 2015 · 2015
Cited alongside, same era.
Layer normalization
Ba, J. L.; Kiros, J. R.; and Hinton, G. E. 2016 · 2016
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, M.; Recht, B.; and Singer, Y. 2016 · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S.; Mudigere, D.; Nocedal, J.; Smelyanskiy, M.; and Tang, P. T. P. 2016 · 2016
Cited alongside, same era.
A closer look at memorization in deep networks
Arpit, D.; Jastrzkebski, S.; Ballas, N.; Krueger, D.; Bengio, E.; Kanwal, M. S.; Maharaj, T.; Fischer, A.; Courville, A.; Bengio, Y.; and Lacoste-Julien, S. 2017 · 2017
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Bartlett, P. L.; Foster, D. J.; and Telgarsky, M. J. 2017 · 2017
Cited alongside, same era.
Retraining: A simple way to improve the ensemble accuracy of deep neural networks for image classification
Zhao, K.; Matsukawa, T.; and Suzuki, E. 2018 · 2018
Later among the works it cites.
Entropy-SGD: Biasing gradient descent into wide valleys
Chaudhari, P.; Choromanska, A.; Soatto, S.; LeCun, Y.; Baldassi, C.; Borgs, C.; Chayes, J.; Sagun, L.; and Zecchina, R. 2019 · 2019
Later among the works it cites.
Benchmarking Neural Network Robustness to Common Corruptions and Perturbations
Hendrycks, D.; and Dietterich, T. 2019 · 2019
Later among the works it cites.
Transfusion: Understanding transfer learning for medical imaging
Raghu, M.; Zhang, C.; Kleinberg, J.; and Bengio, S. 2019 · 2019
Later among the works it cites.
RIFLE: Backpropagation in Depth for Deep Transfer Learning through Re-Initializing the Fully-connected LayEr
Li, X.; Xiong, H.; An, H.; Xu, C.-Z.; and Dou, D. 2020 · 2020
Later among the works it cites.
What Do Neural Networks Learn When Trained With Random Labels?
Maennel, H.; Alabdulmohsin, I.; Tolstikhin, I.; Baldock, R. J.; Bousquet, O.; Gelly, S.; and Keysers, D. 2020 · 2020
Later among the works it cites.
What is being transferred in transfer learning?
Neyshabur, B.; Sedghi, H.; and Zhang, C. 2020 · 2020
Later among the works it cites.
Deep Learning Through the Lens of Example Difficulty
Baldock, R. J.; Maennel, H.; and Neyshabur, B. 2021 · 2021
Closest in time.
Sharpness-Aware Minimization for Efficiently Improving Generalization
Foret, P.; Kleiner, A.; Mobahi, H.; and Neyshabur, B. 2021 · 2021
Closest in time.
Knowledge Evolution in Neural Networks
Taha, A.; Shrivastava, A.; and Davis, L. 2021 · 2021
Closest in time.