Fetching the paper…
Reading the bibliography…
Understanding the implicit regularization (or implicit bias) of gradient descent has recently been a very active research area.
In search of the real inductive bias: On the role of implicit regularization in deep learning
B. Neyshabur, R. Tomioka, and N. Srebro · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
S. Shalev-Shwartz and S. Ben-David · 2014
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2016
Earlier work this paper cites.
Exploring generalization in deep learning
B. Neyshabur, S. Bhojanapalli, D. McAllester, and N. Srebro · 2017
Earlier work this paper cites.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
S. S. Du, W. Hu, and J. D. Lee · 2018
Earlier work this paper cites.
Implicit regularization in matrix factorization
S. Gunasekar, B. Woodworth, S. Bhojanapalli, B. Neyshabur, and N. Srebro · 2018
Earlier work this paper cites.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Y. Li, T. Ma, and H. Zhang · 2018
Earlier work this paper cites.
Implicit regularization in nonconvex statistical estimation: Gradient descent converges linearly for phase retrieval and matrix completion
C. Ma, K. Wang, Y. Chi, and Y. Chen · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
D. Soudry, E. Hoffer, M. S. Nacson, S. Gunasekar, and N. Srebro · 2018
Earlier work this paper cites.
Connecting optimization and regularization paths
A. Suggala, A. Prasad, and P. K. Ravikumar · 2018
Earlier work this paper cites.
When will gradient methods converge to max-margin classifier under relu models?
T. Xu, Y. Zhou, K. Ji, and Y. Liang · 2018
Cited alongside, same era.
Implicit regularization in deep matrix factorization
S. Arora, N. Cohen, W. Hu, and Y. Luo · 2019
Cited alongside, same era.
Implicit regularization of discrete gradient dynamics in linear neural networks
G. Gidel, F. Bach, and S. Lacoste-Julien · 2019
Cited alongside, same era.
Gradient descent maximizes the margin of homogeneous neural networks
K. Lyu and J. Li · 2019
Cited alongside, same era.
Convergence of gradient descent on separable data
M. S. Nacson, J. Lee, S. Gunasekar, P. H. P. Savarese, N. Srebro, and D. Soudry · 2019
Cited alongside, same era.
Implicit regularization in matrix sensing: A geometric view leads to stronger results
A. Eftekhari and K. Zygalakis · 2020
Closest in time.
Gradient descent follows the regularization path for general losses
Z. Ji, M. Dudík, R. E. Schapire, and M. Telgarsky · 2020
Closest in time.
Implicit bias of gradient descent for mean squared error regression with wide neural networks
H. Jin and G. Montúfar · 2020
Closest in time.
Z. Li, Y. Luo, and K. Lyu · 2020
Closest in time.
Implicit bias in deep linear classification: Initialization scale vs training accuracy
E. Moroshko, B. E. Woodworth, S. Gunasekar, J. D. Lee, N. Srebro, and D. Soudry · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Overparameterized nonlinear learning: Gradient descent takes the shortest path?
S. Oymak and M. Soltanolkotabi · 2019
Cited alongside, same era.
Gradient dynamics of shallow univariate relu networks
F. Williams, M. Trager, D. Panozzo, C. Silva, D. Zorin, and J. Bruna · 2019
Cited alongside, same era.
On implicit regularization: Morse functions and applications to matrix factorization
M. A. Belabbas · 2020
Cited alongside, same era.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
L. Chizat and F. Bach · 2020
Cited alongside, same era.
Can implicit bias explain generalization? stochastic convex optimization as a case study
A. Dauber, M. Feder, T. Koren, and R. Livni · 2020
Cited alongside, same era.
Characterizing implicit bias in terms of optimization geometry
S. Gunasekar, J. Lee, D. Soudry, and N. Srebro
Cited in the paper.
Implicit bias of gradient descent on linear convolutional networks
S. Gunasekar, J. D. Lee, D. Soudry, and N. Srebro
Cited in the paper.
N. Razin and N. Cohen · 2020
Closest in time.
Gradient methods never overfit on separable data
O. Shamir · 2020
Closest in time.
Kernel and rich regimes in overparametrized models
B. Woodworth, S. Gunasekar, J. D. Lee, E. Moroshko, P. Savarese, I. Golan, D. Soudry, and N. Srebro · 2020
Closest in time.
Learning a single neuron with gradient methods
G. Yehudai and O. Shamir · 2020
Closest in time.
A unifying view on implicit bias in training linear neural networks
C. Yun, S. Krishnan, and H. Mobahi · 2020
Closest in time.