Fetching the paper…
Reading the bibliography…
Modern deep neural networks are highly over-parameterized compared to the data on which they are trained, yet they often generalize remarkably well.
Sgd on neural networks learns functions of increasing complexity
Nakkiran, P., Kaplun, G., Kalimeris, D., Yang, T., Edelman, B. L., Zhang, F., and Barak, B · 1905
Earlier work this paper cites.
Deep double descent: Where bigger models and more data hurt
Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B., and Sutskever, I · 1912
Earlier work this paper cites.
Praktische verfahren der gleichungsauflösung
Mises, R. and Pollaczek-Geiringer, H · 1929
Earlier work this paper cites.
Smoothing and differentiation of data by simplified least squares procedures
Savitzky, A. and Golay, M. J · 1964
Earlier work this paper cites.
A formal theory of inductive inference. part i
Solomonoff, R. J · 1964
Earlier work this paper cites.
A method for unconstrained convex minimization problem with the rate of convergence o (1/kˆ 2)
Nesterov, Y · 1983
Earlier work this paper cites.
On the limited memory bfgs method for large scale optimization
Liu, D. C. and Nocedal, J · 1989
Earlier work this paper cites.
Neural networks and the bias/variance dilemma
Geman, S., Bienenstock, E., and Doursat, R · 1992
Earlier work this paper cites.
Linear Algebra
Friedberg, S., Insel, A., and Spence, L · 2003
Earlier work this paper cites.
Reducing the time complexity of the derandomized evolution strategy with covariance matrix adaptation (cma-es)
Hansen, N., Müller, S. D., and Koumoutsakos, P · 2003
Earlier work this paper cites.
Clustering methods
Rokach, L. and Maimon, O · 2005
Earlier work this paper cites.
The effective rank: A measure of effective dimensionality
Roy, O. and Vetterli, M · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Eigenvalues and singular values of products of rectangular gaussian random matrices
Burda, Z., Jarosz, A., Livan, G., Nowak, M. A., and Swiech, A · 2010
Earlier work this paper cites.
Kernel analysis of deep networks
Montavon, G., Braun, M. L., and Müller, K.-R · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Plancherel–rotach formulae for average characteristic polynomials of products of ginibre random matrices and the fuss–catalan distribution
Neuschel, T · 2014
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural network
Saxe, A. M., Mcclelland, J. L., and Ganguli, S · 2014
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Goodfellow, I. J., Vinyals, O., and Saxe, A. M · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Skip-thought vectors
Kiros, R., Zhu, Y., Salakhutdinov, R., Zemel, R. S., Torralba, A., Urtasun, R., and Fidler, S · 2015
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Neyshabur, B., Tomioka, R., and Srebro, N · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al · 2015
Earlier work this paper cites.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2015
Earlier work this paper cites.
The power of depth for feedforward neural networks
Eldan, R. and Shamir, O · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Hendrycks, D. and Gimpel, K · 2016
Cited alongside, same era.
Low-rank solutions of linear matrix equations via procrustes flow
Tu, S., Boczar, R., Simchowitz, M., Soltanolkotabi, M., and Recht, B · 2016
Cited alongside, same era.
A closer look at memorization in deep networks
Arpit, D., Jastrzębski, S., Ballas, N., Krueger, D., Bengio, E., Kanwal, M. S., Maharaj, T., Fischer, A., Courville, A., Bengio, Y., et al · 2017
Cited alongside, same era.
Implicit regularization in matrix factorization
Gunasekar, S., Woodworth, B. E., Bhojanapalli, S., Neyshabur, B., and Srebro, N · 2017
Cited alongside, same era.
Deep learning scaling is predictable, empirically
Hestness, J., Narang, S., Ardalani, N., Diamos, G., Jun, H., Kianinejad, H., Patwary, M., Ali, M., Yang, Y., and Zhou, Y · 2017
Cited alongside, same era.
Neural reparameterization improves structural optimization
Hoyer, S., Sohl-Dickstein, J., and Greydanus, S · 2019
Later among the works it cites.
Convergence of gradient descent on separable data
Nacson, M. S., Lee, J., Gunasekar, S., Savarese, P. H. P., Srebro, N., and Soudry, D · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Later among the works it cites.
Stable rank normalization for improved generalization in neural networks and gans
Sanyal, A., Torr, P. H., and Dokania, P. K · 2019
Later among the works it cites.
Deep learning generalizes because the parameter-function map is biased towards simple functions
Valle-Perez, G., Camargo, C. Q., and Louis, A. A · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Building a regular decision boundary with deep networks
Oyallon, E · 2017
Cited alongside, same era.
Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice
Pennington, J., Schoenholz, S. S., and Ganguli, S · 2017
Cited alongside, same era.
Aggregated residual transformations for deep neural networks
Xie, S., Girshick, R., Dollár, P., Tu, Z., and He, K · 2017
Cited alongside, same era.
Spectral norm regularization for improving the generalizability of deep learning
Yoshida, Y. and Miyato, T · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2017
Cited alongside, same era.
Neural architecture search with reinforcement learning
Zoph, B. and Le, Q. V · 2017
Cited alongside, same era.
On the optimization of deep networks: Implicit acceleration by overparameterization
Arora, S., Cohen, N., and Hazan, E · 2018
Cited alongside, same era.
Later among the works it cites.
Global convergence of adaptive gradient methods for an over-parameterized neural network
Wu, X., Du, S. S., and Ward, R · 2019
Later among the works it cites.
A fine-grained spectral perspective on neural networks
Yang, G. and Salman, H · 2019
Later among the works it cites.
Fast convergence of natural gradient descent for over-parameterized neural networks
Zhang, G., Martens, J., and Grosse, R. B · 2019
Later among the works it cites.
High-dimensional dynamics of generalization error in neural networks
Advani, M. S., Saxe, A. M., and Sompolinsky, H · 2020
Later among the works it cites.
Why bigger is not always better: on finite and infinite neural networks
Aitchison, L · 2020
Later among the works it cites.
Benign overfitting in linear regression
Bartlett, P. L., Long, P. M., Lugosi, G., and Tsigler, A · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Later among the works it cites.
Expandnets: Linear over-parameterization to train compact convolutional networks
Guo, S., Alvarez, J. M., and Salzmann, M · 2020
Later among the works it cites.
Directional convergence and alignment in deep learning
Ji, Z. and Telgarsky, M · 2020
Later among the works it cites.
Implicit rank-minimizing autoencoder
Jing, L., Zbontar, J., et al · 2020
Later among the works it cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Later among the works it cites.
Li, Z., Luo, Y., and Lyu, K · 2020
Later among the works it cites.
Do deeper convolutional networks perform better?
Nichani, E., Radhakrishnan, A., and Uhler, C · 2020
Later among the works it cites.
Gradient starvation: A learning proclivity in neural networks
Pezeshki, M., Kaba, S.-O., Bengio, Y., Courville, A., Precup, D., and Lajoie, G · 2020
Later among the works it cites.
Implicit regularization in deep learning may not be explainable by norms
Razin, N. and Cohen, N · 2020
Later among the works it cites.
The pitfalls of simplicity bias in neural networks
Shah, H., Tamuly, K., Raghunathan, A., Jain, P., and Netrapalli, P · 2020
Later among the works it cites.
Implicit neural representations with periodic activation functions
Sitzmann, V., Martel, J., Bergman, A., Lindell, D., and Wetzstein, G · 2020
Later among the works it cites.
Deep kernel processes
Aitchison, L., Yang, A., and Ober, S. W · 2021
Closest in time.
Implicit regularization via neural feature alignment
Baratin, A., George, T., Laurent, C., Hjelm, R. D., Lajoie, G., Vincent, P., and Lacoste-Julien, S · 2021
Closest in time.
Deep learning: a statistical viewpoint
Bartlett, P. L., Montanari, A., and Rakhlin, A · 2021
Closest in time.
Are wider nets better given the same number of parameters?
Golubeva, A., Neyshabur, B., and Gur-Ari, G · 2021
Closest in time.
Asymptotics of representation learning in finite bayesian neural networks
Zavatone-Veth, J. A., Canatar, A., Ruben, B., and Pehlevan, C · 2021
Closest in time.