Fetching the paper…
Reading the bibliography…
Highly overparametrized neural networks can display curiously strong generalization performance - a phenomenon that has recently garnered a wealth of theoretical and empirical research in order to better understand it.
Modeling by shortest data description
Rissanen, J · 1978
Earlier work this paper cites.
Present position and potential developments: Some personal views statistical theory the prequential approach
Dawid, A. P · 1984
Earlier work this paper cites.
Stochastic complexity in statistical inquiry
Rissanen, J · 1989
Earlier work this paper cites.
Neural networks and the bias/variance dilemma
Geman, S., Bienenstock, E., and Doursat, R · 1992
Earlier work this paper cites.
The mnist database of handwritten digits
LeCun, Y · 1998
Earlier work this paper cites.
A unified bias-variance decomposition
Domingos, P · 2000
Earlier work this paper cites.
Information theory, inference and learning algorithms
Mac Kay, D. J · 2003
Earlier work this paper cites.
The Minimum Description Length Principle
Grünwald, P. D · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2016
Cited alongside, same era.
Deep learning scaling is predictable, empirically
Hestness, J., Narang, S., Ardalani, N., Diamos, G., Jun, H., Kianinejad, H., Patwary, M., Ali, M., Yang, Y., and Zhou, Y · 2017
Later among the works it cites.
Stochastic gradient descent as approximate bayesian inference
Mandt, S., Hoffman, M. D., and Blei, D. M · 2017
Later among the works it cites.
The description length of deep learning models
Blier, L. and Ollivier, Y · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Later among the works it cites.
A jamming transition from under-to over-parametrization affects loss landscape and generalization
Spigler, S., Geiger, M., d’Ascoli, S., Sagun, L., Biroli, G., and Wyart, M · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zagoruyko, S. and Komodakis, N · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2016
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks
Advani, M. S. and Saxe, A. M · 2017
Cited alongside, same era.
Emnist: Extending mnist to handwritten letters
Cohen, G., Afshar, S., Tapson, J., and Van Schaik, A · 2017
Cited alongside, same era.
On calibration of modern neural networks
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q · 2017
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Allen-Zhu, Z., Li, Y., and Liang, Y · 2019
Later among the works it cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Belkin, M., Hsu, D., Ma, S., and Mandal, S · 2019
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B., and Sutskever, I · 2019
Later among the works it cites.
Nas-bench-101: Towards reproducible neural architecture search
Ying, C., Klein, A., Real, E., Christiansen, E., Murphy, K., and Hutter, F · 2019
Later among the works it cites.
Towards robust image classification using sequential attention models
Zoran, D., Chrzanowski, M., Huang, P.-S., Gowal, S., Mott, A., and Kohl, P · 2019
Later among the works it cites.