Fetching the paper…
Reading the bibliography…
In deep learning, it is common to use more network parameters than training points.
Principal component analysis
S. Wold, K. Esbensen, and P. Geladi · 1987
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
P. Baldi and K. Hornik · 1989
Earlier work this paper cites.
The MNIST database of handwritten digits
Y. LeCun, C. Cortes, and C. Burges · 1998
Earlier work this paper cites.
On early stopping in gradient descent learning
Y. Yao, L. Rosasco, and A. Caponnetto · 2007
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Matrix analysis
R. A. Horn and C. R. Johnson · 2012
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
B. Neyshabur, R. Tomioka, and N. Srebro · 2015
Earlier work this paper cites.
Stable low-rank matrix recovery via null space properties
M. Kabanava, R. Kueng, H. Rauhut, and U. Terstiege · 2016
Earlier work this paper cites.
Deep learning without poor local minima
K. Kawaguchi · 2016
Earlier work this paper cites.
Implicit regularization in matrix factorization
S. Gunasekar, B. E. Woodworth, S. Bhojanapalli, B. Neyshabur, and N. Srebro · 2017
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2017
Cited alongside, same era.
Geometry of optimization and implicit regularization in deep learning
B. Neyshabur, R. Tomioka, R. Salakhutdinov, and N. Srebro · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Cited alongside, same era.
On the optimization of deep networks: Implicit acceleration by overparameterization
S. Arora, N. Cohen, and E. Hazan · 2018
Cited alongside, same era.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Y. Li, T. Ma, and H. Zhang · 2018
Later among the works it cites.
The implicit bias of gradient descent on separable data
D. Soudry, E. Hoffer, M. S. Nacson, S. Gunasekar, and N. Srebro · 2018
Later among the works it cites.
Deep image prior
D. Ulyanov, A. Vedaldi, and V. Lempitsky · 2018
Later among the works it cites.
Implicit regularization in deep matrix factorization
S. Arora, N. Cohen, W. Hu, and Y. Luo · 2019
Later among the works it cites.
Implicit regularization of discrete gradient dynamics in linear neural networks
G. Gidel, F. Bach, and S. Lacoste-Julien · 2019
Later among the works it cites.
Deep decoder: Concise image representations from untrained non-convolutional networks
R. Heckel and P. Hand · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gradient descent with identity initialization efficiently learns positive definite linear transformations by deep residual networks
P. Bartlett, D. Helmbold, and P. Long · 2018
Cited alongside, same era.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
S. S. Du, W. Hu, and J. Lee · 2018
Cited alongside, same era.
Implicit bias of gradient descent on linear convolutional networks
S. Gunasekar, J. D. Lee, D. Soudry, and N. Srebro · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Cited alongside, same era.
Learning deep linear neural networks: Riemannian gradient flows and convergence to global minimizers
B. Bah, H. Rauhut, U. Terstiege, and M. Westdickenberg
Cited in the paper.
Later among the works it cites.
Low-rank regularization and solution uniqueness in over-parameterized matrix sensing
K. Geyer, A. Kyrillidis, and A. Kalev · 2020
Closest in time.
The implicit bias of depth: How incremental learning drives generalization
D. Gissin, S. Shalev-Shwartz, and A. Daniely · 2020
Closest in time.
Implicit regularization in deep learning may not be explainable by norms
N. Razin and N. Cohen · 2020
Closest in time.