Fetching the paper…
Reading the bibliography…
This study demonstrates that double descent can be mitigated by adding a dropout layer adjacent to the fully connected linear layer.
1903
Earlier work this paper cites.
Srivastava N, Hinton G, Krizhevsky A, et al (2014) Dropout: A simple way to prevent neural networks from overfitting. Journal of Machine Learning Research 15(56):1929–1958. URL http://jmlr.org/papers/v15/srivastava14a.html
1958
Earlier work this paper cites.
Hemmerle WJ (1975) An explicit solution for generalized ridge regression. Technometrics 17(3):309–314. URL http://www.jstor.org/stable/1268066
1975
Earlier work this paper cites.
Strawderman WE (1978) Minimax adaptive generalized ridge regression estimators. Journal of the American Statistical Association 73(363):623–627. URL http://www.jstor.org/stable/2286612
1978
Earlier work this paper cites.
Casella G (1980) Minimax Ridge Regression Estimation. The Annals of Statistics 8(5):1036 – 1056. 10.1214/aos/1176345141 , URL https://doi.org/10.1214/aos/1176345141
1980
Earlier work this paper cites.
Hua TA, Gunst RF (1983) Generalized ridge regression: a note on negative ridge parameters. Communications in Statistics - Theory and Methods 12(1):37–45. 10.1080/03610928308828440 , URL https://doi.org/10.1080/03610928308828440 , https://doi.org/10.1080/03610928308828440
1983
Earlier work this paper cites.
Cun YL, Kanter I, Solla SA (1991) Eigenvalues of covariance matrices: Application to neural-network learning. Phys Rev Lett 66:2396–2399. 10.1103/PhysRevLett.66.2396 , URL https://link.aps.org/doi/10.1103/PhysRevLett.66.2396
1991
Earlier work this paper cites.
Opper M, Kinzel W (1996) Statistical Mechanics of Generalization, Springer New York, New York, NY, pp 151–209. 10.1007/978-1-4612-0723-8_5 , URL https://doi.org/10.1007/978-1-4612-0723-8_5
1996
Earlier work this paper cites.
Hoerl AE, Kennard RW (2000) Ridge regression: Biased estimation for nonorthogonal problems. Technometrics 42(1):80–86. URL http://www.jstor.org/stable/1271436
2000
Earlier work this paper cites.
Maruyama Y, Strawderman WE (2005) A new class of generalized Bayes minimax ridge regression estimators. The Annals of Statistics 33(4):1753 – 1770. 10.1214/009053605000000327 , URL https://doi.org/10.1214/009053605000000327
2005
Earlier work this paper cites.
Jiang T (2004) The limiting distributions of eigenvalues of sample correlation matrices. Sankhyā: The Indian Journal of Statistics (2003-2007) 66(1):35–48. URL http://www.jstor.org/stable/25053330
2007
Earlier work this paper cites.
Rahimi A, Recht B (2007) Random features for large-scale kernel machines. In: Platt J, Koller D, Singer Y, et al (eds) Advances in Neural Information Processing Systems, vol 20. Curran Associates, Inc., URL https://proceedings.neurips.cc/paper/2007/file/013a006f03dbc5392effeb8f18fda755-Paper.pdf
2007
Earlier work this paper cites.
Hastie T, Tibshirani R, Friedman J (2009) The Elements of Statistical Learning: Data Mining, Inference, and Prediction., 2nd edn. Springer New York, NY
2009
Earlier work this paper cites.
Krizhevsky A (2009) Learning multiple layers of features from tiny images
2009
Earlier work this paper cites.
2012
Earlier work this paper cites.
Krizhevsky A, Sutskever I, Hinton GE (2012) Imagenet classification with deep convolutional neural networks. In: Pereira F, Burges C, Bottou L, et al (eds) Advances in Neural Information Processing Systems, vol 25. Curran Associates, Inc., URL https://proceedings.neurips.cc/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf
2012
Earlier work this paper cites.
Ba LJ, Frey B (2013) Adaptive dropout for training deep neural networks. In: Proceedings of the 26th International Conference on Neural Information Processing Systems - Volume 2. Curran Associates Inc., Red Hook, NY, USA, NIPS’13, pp 3084–3092
2013
Cited alongside, same era.
2013
Cited alongside, same era.
Wang S, Manning C (2013) Fast dropout training. In: Dasgupta S, McAllester D (eds) Proceedings of the 30th International Conference on Machine Learning, Proceedings of Machine Learning Research, vol 28. PMLR, Atlanta, Georgia, USA, pp 118–126, URL https://proceedings.mlr.press/v28/wang13a.html
2013
Cited alongside, same era.
Ishwaran H, Rao JS (2014) Geometry and properties of generalized ridge regression in high dimensions
2014
Cited alongside, same era.
Page D (2018) How to train your resnet 4: Architecture. URL https://myrtle.ai/how-to-train-your-resnet-4-architecture/
2018
Later among the works it cites.
Saito K, Ushiku Y, Harada T, et al (2018) Adversarial dropout regularization. In: International Conference on Learning Representations, URL https://openreview.net/forum?id=HJIoJWZCZ
2018
Later among the works it cites.
Belkin M, Hsu D, Ma S, et al (2019) Reconciling modern machine-learning practice and the classical bias–variance trade-off. Proceedings of the National Academy of Sciences 116(32):15849–15854. 10.1073/pnas.1903070116 , URL https://www.pnas.org/doi/abs/10.1073/pnas.1903070116 , https://www.pnas.org/doi/pdf/10.1073/pnas.1903070116
2019
Later among the works it cites.
Chen G, Chen P, Shi Y, et al (2019) Rethinking the usage of batch normalization and dropout in the training of deep neural networks. arXiv preprint arXiv:190505928
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Toshev A, Szegedy C (2014) Deeppose: Human pose estimation via deep neural networks. In: 2014 IEEE Conference on Computer Vision and Pattern Recognition, pp 1653–1660, 10.1109/CVPR.2014.214
2014
Cited alongside, same era.
Helmbold DP, Long PM (2015) On the inductive bias of dropout. Journal of Machine Learning Research 16(105):3403–3454. URL http://jmlr.org/papers/v16/helmbold15a.html
2015
Cited alongside, same era.
2015
Cited alongside, same era.
Kingma DP, Salimans T, Welling M (2015) Variational dropout and the local reparameterization trick. In: Cortes C, Lawrence N, Lee D, et al (eds) Advances in Neural Information Processing Systems, vol 28. Curran Associates, Inc., URL https://proceedings.neurips.cc/paper/2015/file/bc7316929fe1545bf0b98d114ee3ecb8-Paper.pdf
2015
Cited alongside, same era.
Gal Y, Ghahramani Z (2016) Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In: Balcan MF, Weinberger KQ (eds) Proceedings of The 33rd International Conference on Machine Learning, Proceedings of Machine Learning Research, vol 48. PMLR, New York, New York, USA, pp 1050–1059, URL https://proceedings.mlr.press/v48/gal16.html
2016
Cited alongside, same era.
Gao W, Zhou ZH (2016) Dropout rademacher complexity of deep neural networks. Science China Information Sciences 59(7):072104. 10.1007/s11432-015-5470-z , URL https://doi.org/10.1007/s11432-015-5470-z
2016
Cited alongside, same era.
Li Z, Gong B, Yang T (2016) Improved dropout for shallow and deep learning. In: Proceedings of the 30th International Conference on Neural Information Processing Systems. Curran Associates Inc., Red Hook, NY, USA, NIPS’16, pp 2531–2539
2016
Cited alongside, same era.
Arpit D, Jastrzkebski S, Ballas N, et al (2017) A closer look at memorization in deep networks. In: Precup D, Teh YW (eds) Proceedings of the 34th International Conference on Machine Learning, Proceedings of Machine Learning Research, vol 70. PMLR, pp 233–242, URL https://proceedings.mlr.press/v70/arpit17a.html
2017
Cited alongside, same era.
Li X, Chen S, Hu X, et al (2019) Understanding the disharmony between dropout and batch normalization by variance shift. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 2682–2690
2019
Later among the works it cites.
Wainwright MJ (2019) High-Dimensional Statistics: A Non-Asymptotic Viewpoint. Cambridge Series in Statistical and Probabilistic Mathematics, Cambridge University Press, 10.1017/9781108627771
2019
Later among the works it cites.
Advani MS, Saxe AM, Sompolinsky H (2020) High-dimensional dynamics of generalization error in neural networks. Neural Networks 132:428–446. https://doi.org/10.1016/j.neunet.2020.08.022 , URL https://www.sciencedirect.com/science/article/pii/S0893608020303117
2020
Later among the works it cites.
Belkin M, Hsu D, Xu J (2020) Two models of double descent for weak features. SIAM Journal on Mathematics of Data Science 2(4):1167–1180. 10.1137/20M1336072 , URL https://doi.org/10.1137/20M1336072 , https://doi.org/10.1137/20M1336072
2020
Later among the works it cites.
Garbin C, Zhu X, Marques O (2020) Dropout vs. batch normalization: an empirical study of their impact to deep learning. Multimedia Tools and Applications 79:12777–12815
2020
Later among the works it cites.
LeJeune D, Javadi H, Baraniuk R (2020) The implicit regularization of ordinary least squares ensembles. In: Chiappa S, Calandra R (eds) Proceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics, Proceedings of Machine Learning Research, vol 108. PMLR, pp 3525–3535, URL https://proceedings.mlr.press/v108/lejeune20b.html
2020
Later among the works it cites.
Wei C, Kakade S, Ma T (2020) The implicit and explicit regularization effects of dropout. In: III HD, Singh A (eds) Proceedings of the 37th International Conference on Machine Learning, Proceedings of Machine Learning Research, vol 119. PMLR, pp 10181–10192, URL https://proceedings.mlr.press/v119/wei20d.html
2020
Later among the works it cites.
Wu D, Xu J (2020) On the optimal weighted \ell_2 regularization in overparameterized linear regression. In: Larochelle H, Ranzato M, Hadsell R, et al (eds) Advances in Neural Information Processing Systems, vol 33. Curran Associates, Inc., pp 10112–10123, URL https://proceedings.neurips.cc/paper/2020/file/72e6d3238361fe70f22fb0ac624a7072-Paper.pdf
2020
Later among the works it cites.
Chen L, Min Y, Belkin M, et al (2021) Multiple descent: Design your own generalization curve. In: Ranzato M, Beygelzimer A, Dauphin Y, et al (eds) Advances in Neural Information Processing Systems, vol 34. Curran Associates, Inc., pp 8898–8912, URL https://proceedings.neurips.cc/paper/2021/file/4ae67a7dd7e491f8fb6f9ea0cf25dfdb-Paper.pdf
2021
Later among the works it cites.
Heckel R, Yilmaz FF (2021) Early stopping in deep networks: Double descent and how to eliminate it. In: International Conference on Learning Representations, URL https://openreview.net/forum?id=tlV90jvZbw
2021
Later among the works it cites.
Nakkiran P, Kaplun G, Bansal Y, et al (2021a) Deep double descent: where bigger models and more data hurt*. Journal of Statistical Mechanics: Theory and Experiment 2021(12):124003. 10.1088/1742-5468/ac3a74 , URL https://dx.doi.org/10.1088/1742-5468/ac3a74
2021
Later among the works it cites.