Fetching the paper…
Reading the bibliography…
We study the implicit regularization of optimization methods for linear models interpolating the training data in the under-parametrized and over-parametrized regimes.
A method for the solution of certain non-linear problems in least squares
Kenneth Levenberg · 1944
Earlier work this paper cites.
An algorithm for least-squares estimation of nonlinear parameters
Donald W Marquardt · 1963
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Kernel methods for pattern analysis
John Shawe-Taylor, Nello Cristianini, et al · 2004
Earlier work this paper cites.
Ellipsoidal machines
Pannagadatta K Shivaswamy and Tony Jebara · 2007
Earlier work this paper cites.
On the complexity of linear prediction: Risk bounds, margin bounds, and regularization
Sham M Kakade, Karthik Sridharan, and Ambuj Tewari · 2009
Earlier work this paper cites.
Maximum relative margin and data-dependent regularization
Pannagadatta K Shivaswamy and Tony Jebara · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Sublinear optimization for machine learning
Kenneth L Clarkson, Elad Hazan, and David P Woodruff · 2012
Earlier work this paper cites.
Do we need hundreds of classifiers to solve real world classification problems?
Manuel Fernández-Delgado, Eva Cernadas, Senén Barro, and Dinani Amorim · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
UCI machine learning repository, 2017
Dheeru Dua and Casey Graff · 2017
Cited alongside, same era.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Cited alongside, same era.
Improving generalization performance by switching from adam to sgd
Nitish Shirish Keskar and Richard Socher · 2017
Cited alongside, same era.
Training deep networks without learning rates through coin betting
Francesco Orabona and Tatiana Tommasi · 2017
Cited alongside, same era.
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 2019
Later among the works it cites.
Harnessing the power of infinitely wide deep nets on small-data tasks
Sanjeev Arora, Simon S Du, Zhiyuan Li, Ruslan Salakhutdinov, Ruosong Wang, and Dingli Yu · 2019
Later among the works it cites.
Benign overfitting in linear regression
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler · 2019
Later among the works it cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Later among the works it cites.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Cited alongside, same era.
To understand deep learning we need to understand kernel learning
Mikhail Belkin, Siyuan Ma, and Soumik Mandal · 2018
Cited alongside, same era.
Characterizing implicit bias in terms of optimization geometry
Suriya Gunasekar · 2018
Cited alongside, same era.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason D Lee, Daniel Soudry, and Nati Srebro · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Gradient descent aligns the layers of deep linear networks
Ziwei Ji and Matus Telgarsky · 2018
Cited alongside, same era.
Convergence of gradient descent on separable data
Mor Shpigel Nacson, Jason Lee, Suriya Gunasekar, Pedro HP Savarese, Nathan Srebro, and Daniel Soudry · 2018
Cited alongside, same era.
Later among the works it cites.
Stochastic gradient descent on separable data: Exact convergence with a fixed learning rate
Mor Shpigel Nacson, Nathan Srebro, and Daniel Soudry · 2019
Later among the works it cites.
The effect of network width on stochastic gradient descent and generalization: an empirical study
Daniel S Park, Jascha Sohl-Dickstein, Quoc V Le, and Samuel L Smith · 2019
Later among the works it cites.
The implicit bias of adagrad on separable data
Qian Qian and Xiaoyuan Qian · 2019
Later among the works it cites.
Finite-sample analysis of interpolating linear classifiers in the overparameterized regime
Niladri S Chatterji and Philip M Long · 2020
Closest in time.
BackPACK: Packing more into backprop
Felix Dangel, Frederik Kunstner, and Philipp Hennig · 2020
Closest in time.
Bias of homotopic gradient descent for the hinge loss
Denali Molitor, Deanna Needell, and Rachel Ward · 2020
Closest in time.
Harmless interpolation of noisy data in regression
Vidya Muthukumar, Kailas Vodrahalli, Vignesh Subramanian, and Anant Sahai · 2020
Closest in time.
Implicit regularization in deep learning may not be explainable by norms
Noam Razin and Nadav Cohen · 2020
Closest in time.