Fetching the paper…
Reading the bibliography…
We provide a detailed study on the implicit bias of gradient descent when optimizing loss functions with strictly monotone tails, such as the logistic loss, over separable datasets.
Margin maximizing loss functions
Saharon Rosset, Ji Zhu, and Trevor J Hastie · 2004
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2004
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Sublinear optimization for machine learning
Kenneth L. Clarkson, Elad Hazan, and David P. Woodruff · 2012
Earlier work this paper cites.
Boosting: Foundations and algorithms
Robert E. Schapire and Yoav Freund · 2012
Earlier work this paper cites.
Margins, shrinkage and boosting
Matus Telgarsky · 2013
Cited alongside, same era.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Cited alongside, same era.
The Power of Normalization: Faster Evasion of Saddle Points
Kfir Y. Levy · 2016
Cited alongside, same era.
Wide residual networks
Sergey Zagoruyko and Nikos Komodakis · 2016
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Cited alongside, same era.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro
Cited in the paper.
Characterizing implicit bias in terms of optimization geometry
Suriya Gunasekar, Jason D. Lee, Daniel Soudry, and Nathan Srebro
Cited in the paper.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, , Mor Shpigel Nacson, and Nathan Srebro
Cited in the paper.
The implicit bias of gradient descent on separable data (journal version)
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro
Cited in the paper.
Risk and parameter convergence of logistic regression
Ziwei Ji and Matus Telgarsky · 2018
Closest in time.
Asymptotic solution for a first order ode
Willie Wong · 2018
Closest in time.
Convergence of sgd in learning relu models with separable data
Tengyu Xu, Yi Zhou, Kaiyi Ji, and Yingbin Liang · 2018
Closest in time.
Gradient descent aligns the layers of deep linear networks
Ziwei Ji and Matus Telgarsky · 2019
Closest in time.
Stochastic gradient descent on separable data: Exact convergence with a fixed learning rate
Mor Shpigel Nacson, Nathan Srebro, and Daniel Soudry · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…