Fetching the paper…
Reading the bibliography…
Controlling the parameters' norm often yields good generalisation when training neural networks.
A simple weight decay can improve generalization
Anders Krogh and John Hertz · 1991
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Andrew R Barron · 1993
Earlier work this paper cites.
For valid generalization the size of the weights is more important than the size of the network
Peter Bartlett · 1996
Earlier work this paper cites.
Atomic decomposition by basis pursuit
Scott Shaobing Chen, David L Donoho, and Michael A Saunders · 2001
Earlier work this paper cites.
Bounds on rates of variable-basis and neural-network approximation
Vera Kurková and Marcello Sanguineti · 2001
Earlier work this paper cites.
Convex neural networks
Yoshua Bengio, Nicolas Roux, Pascal Vincent, Olivier Delalleau, and Patrice Marcotte · 2005
Earlier work this paper cites.
Complexity measures for neural networks with general activation functions using path-based norms
Zhong Li, Chao Ma, and Lei Wu · 2009
Earlier work this paper cites.
Sparse autoencoder
Andrew Ng · 2011
Earlier work this paper cites.
Zhiyuan Li, Yuping Luo, and Kaifeng Lyu · 2012
Earlier work this paper cites.
Towards a mathematical theory of super-resolution
Emmanuel J Candès and Carlos Fernandez-Granda · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
Super-resolution of point sources via convex programming
Carlos Fernandez-Granda · 2016
Earlier work this paper cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2017
Cited alongside, same era.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Cited alongside, same era.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Cited alongside, same era.
The implicit bias of gradient descent on nonseparable data
Ziwei Ji and Matus Telgarsky · 2019
Banach space representer theorems for neural networks and ridge splines
Rahul Parhi and Robert D Nowak · 2021
Later among the works it cites.
Implicit regularization in relu networks with the square loss
Gal Vardi and Ohad Shamir · 2021
Later among the works it cites.
The hidden convex optimization landscape of regularized two-layer relu networks: an exact characterization of optimal solutions
Yifei Wang, Jonathan Lacotte, and Mert Pilanci · 2021
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Later among the works it cites.
Gradient flow dynamics of shallow relu networks for square loss and orthogonal inputs
Etienne Boursier, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2019
Cited alongside, same era.
A function space view of bounded norm infinite width relu nets: The multivariate case
Greg Ongie, Rebecca Willett, Daniel Soudry, and Nathan Srebro · 2019
Cited alongside, same era.
Support localization and the fisher metric for off-the-grid sparse regularization
Clarice Poon, Nicolas Keriven, and Gabriel Peyré · 2019
Cited alongside, same era.
How do infinite width bounded norm networks look in function space?
Pedro Savarese, Itay Evron, Daniel Soudry, and Nathan Srebro · 2019
Cited alongside, same era.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lenaic Chizat and Francis Bach · 2020
Cited alongside, same era.
Implicit regularization in deep learning may not be explainable by norms
Noam Razin and Nadav Cohen · 2020
Cited alongside, same era.
Sparsest piecewise-linear regression of one-dimensional data
Thomas Debarre, Quentin Denoyelle, Michael Unser, and Julien Fageot · 2022
Later among the works it cites.
Implicit bias of the step size in linear diagonal neural networks
Mor Shpigel Nacson, Kavya Ravichandran, Nathan Srebro, and Daniel Soudry · 2022
Later among the works it cites.
On the effective number of linear regions in shallow univariate relu networks: Convergence guarantees and implicit bias
Itay Safran, Gal Vardi, and Jason D Lee · 2022
Later among the works it cites.
Intrinsic dimensionality and generalization properties of the ℛ \mathcal{R} -norm inductive bias
Clayton Sanford, Navid Ardeshir, and Daniel Hsu · 2022
Later among the works it cites.
Mean-field analysis of piecewise linear solutions for wide relu networks
Alexander Shevchenko, Vyacheslav Kungurtsev, and Marco Mondelli · 2022
Later among the works it cites.
Regression as classification: Influence of task formulation on neural network features
Lawrence Stewart, Francis Bach, Quentin Berthet, and Jean-Philippe Vert · 2022
Later among the works it cites.
Learning a neuron by a shallow relu network: Dynamics and implicit bias for correlated inputs
Dmitry Chistikov, Matthias Englert, and Ranko Lazic · 2023
Closest in time.
Deep learning meets sparse regularization: A signal processing perspective
Rahul Parhi and Robert D Nowak · 2023
Closest in time.