Fetching the paper…
Reading the bibliography…
The notion of implicit bias, or implicit regularization, has been suggested as a means to explain the surprising generalization ability of modern-days overparameterized learning algorithms.
The accuracy of the gaussian approximation to the sum of independent variates
A. C. Berry · 1941
Earlier work this paper cites.
On the liapunov limit error in the theory of probability
C.-G. Esseen · 1942
Earlier work this paper cites.
An application of fourier methods to the problem of sharpening the berry-esseen inequality
P. van Beek · 1972
Earlier work this paper cites.
A theory of the learnable
L. G. Valiant · 1984
Earlier work this paper cites.
Learnability and the vapnik-chervonenkis dimension
A. Blumer, A. Ehrenfeucht, D. Haussler, and M. K. Warmuth · 1989
Earlier work this paper cites.
Boosting the margin: A new explanation for the effectiveness of voting methods
R. E. Schapire, Y. Freund, P. Bartlett, W. S. Lee, et al · 1998
Earlier work this paper cites.
Stability and generalization
O. Bousquet and A. Elisseeff · 2002
Earlier work this paper cites.
Boosting with the l 2 loss: regression and classification
P. Bühlmann and B. Yu · 2003
Earlier work this paper cites.
Convex optimization
S. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
Stochastic convex optimization
S. Shalev-Shwartz, O. Shamir, N. Srebro, and K. Sridharan · 2009
Earlier work this paper cites.
Implicit regularization in variational bayesian matrix factorization
S. Nakajima and M. Sugiyama · 2010
Earlier work this paper cites.
Online learning and online convex optimization
S. Shalev-Shwartz et al · 2011
Cited alongside, same era.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
O. Shamir and T. Zhang · 2013
Cited alongside, same era.
Early stopping and non-parametric regression: an optimal data-dependent stopping rule
G. Raskutti, M. J. Wainwright, and B. Yu · 2014
Cited alongside, same era.
Understanding Machine Learning:From Theory to Algorithms
S. Shalev-Shwartz and S. Ben-David · 2014
Cited alongside, same era.
In search of the real inductive bias: On the role of implicit regularization in deep learning
B. Neyshabur, R. Tomioka, and N. Srebro · 2015
Cited alongside, same era.
Generalization of erm in stochastic convex optimization: The dimension strikes back
Implicit regularization in deep learning
B. Neyshabur · 2017
Later among the works it cites.
Geometry of optimization and implicit regularization in deep learning
B. Neyshabur, R. Tomioka, R. Salakhutdinov, and N. Srebro · 2017
Later among the works it cites.
Early stopping for kernel boosting algorithms: A general analysis with localized complexities
Y. Wei, F. Yang, and M. J. Wainwright · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Later among the works it cites.
The implicit bias of gradient descent on separable data
D. Soudry, E. Hoffer, M. S. Nacson, S. Gunasekar, and N. Srebro · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. Feldman · 2016
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
M. Hardt, B. Recht, and Y. Singer · 2016
Cited alongside, same era.
Generalization properties and implicit regularization for multiple passes sgm
J. Lin, R. Camoriano, and L. Rosasco · 2016
Cited alongside, same era.
Implicit regularization in matrix factorization
S. Gunasekar, B. E. Woodworth, S. Bhojanapalli, B. Neyshabur, and N. Srebro · 2017
Cited alongside, same era.
Affine-invariant online optimization and the low-rank experts problem
T. Koren and R. Livni · 2017
Cited alongside, same era.
Characterizing implicit bias in terms of optimization geometry
S. Gunasekar, J. Lee, D. Soudry, and N. Srebro
Cited in the paper.
Implicit bias of gradient descent on linear convolutional networks
S. Gunasekar, J. D. Lee, D. Soudry, and N. Srebro
Cited in the paper.
Connecting optimization and regularization paths
A. Suggala, A. Prasad, and P. K. Ravikumar · 2018
Later among the works it cites.
A continuous-time view of early stopping for least squares regression
A. Ali, J. Z. Kolter, and R. J. Tibshirani · 2019
Later among the works it cites.
Implicit regularization in deep matrix factorization
S. Arora, N. Cohen, W. Hu, and Y. Luo · 2019
Later among the works it cites.
Convergence of gradient descent on separable data
M. S. Nacson, J. D. Lee, S. Gunasekar, P. H. P. Savarese, N. Srebro, and D. Soudry · 2019
Later among the works it cites.
Uniform convergence may be unable to explain generalization in deep learning
V. Nagarajan and J. Z. Kolter · 2019
Later among the works it cites.