Fetching the paper…
Reading the bibliography…
We provide a function space characterization of the inductive bias resulting from minimizing the $\ell_2$ norm of the weights in multi-channel convolutional neural networks with linear activations and empirically test our resulting hypothesis on ReLU networks trained using gradient descent.
A simple weight decay can improve generalization
Anders Krogh and John A. Hertz · 1991
Earlier work this paper cites.
For valid generalization the size of the weights is more important than the size of the network
Peter L. Bartlett · 1996
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L. Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Fast maximum margin matrix factorization for collaborative prediction
Jason D. M. Rennie and Nathan Srebro · 2005
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
MNIST handwritten digit database
Yann LeCun and Corinna Cortes · 2010
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
Peter L. Bartlett, Dylan J. Foster, and Matus Telgarsky · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Earlier work this paper cites.
Risk and parameter convergence of logistic regression
Ziwei Ji and Matus Telgarsky · 2018
Cited alongside, same era.
Gradient descent aligns the layers of deep linear networks
Ziwei Ji and Matus Telgarsky · 2019
Cited alongside, same era.
Lexicographic and depth-sensitive margins in homogeneous and non-homogeneous deep models
Mor Shpigel Nacson, Suriya Gunasekar, Jason D. Lee, Nathan Srebro, and Daniel Soudry · 2019
Cited alongside, same era.
How do infinite width bounded norm networks look in function space?
Pedro Savarese, Itay Evron, Daniel Soudry, and Nathan Srebro · 2019
Cited alongside, same era.
Regularization matters: Generalization and optimization of neural nets v.s. their induced kernel
Colin Wei, Jason D. Lee, Qiang Liu, and Tengyu Ma · 2019
Cited alongside, same era.
A function space view of bounded norm infinite width relu nets: The multivariate case
Greg Ongie, Rebecca Willett, Daniel Soudry, and Nathan Srebro · 2020
Later among the works it cites.
Neural networks are convex regularizers: Exact polynomial-time convex optimization formulations for two-layer networks
Mert Pilanci and Tolga Ergen · 2020
Later among the works it cites.
Implicit regularization in deep learning may not be explainable by norms
Noam Razin and Nadav Cohen · 2020
Later among the works it cites.
Identity crisis: Memorization and generalization under extreme overparameterization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Michael C. Mozer, and Yoram Singer · 2020
Later among the works it cites.
Representation costs of linear neural networks: Analysis and design
Zhen Dai, Mina Karzand, and Nathan Srebro · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Can implicit bias explain generalization? stochastic convex optimization as a case study
Assaf Dauber, Meir Feder, Tomer Koren, and Roi Livni · 2020
Cited alongside, same era.
Implicit convex regularizers of cnn architectures: Convex optimization of two- and three-layer networks in polynomial time
Tolga Ergen and Mert Pilanci · 2020
Cited alongside, same era.
Directional convergence and alignment in deep learning
Ziwei Ji and Matus Telgarsky · 2020
Cited alongside, same era.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2020
Cited alongside, same era.
Characterizing implicit bias in terms of optimization geometry
Suriya Gunasekar, Jason D. Lee, Daniel Soudry, and Nathan Srebro
Cited in the paper.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason D. Lee, Daniel Soudry, and Nati Srebro
Cited in the paper.
Tolga Ergen and Mert Pilanci · 2021
Closest in time.
Towards resolving the implicit bias of gradient descent for matrix factorization: Greedy low-rank learning
Zhiyuan Li, Yuping Luo, and Kaifeng Lyu · 2021
Closest in time.
Vector-output relu neural network problems are copositive programs: Convex analysis of two layer networks and polynomial-time algorithms
Arda Sahiner, Tolga Ergen, John M. Pauly, and Mert Pilanci · 2021
Closest in time.
A unifying view on implicit bias in training linear neural networks
Chulhee Yun, Shankar Krishnan, and Hossein Mobahi · 2021
Closest in time.