Fetching the paper…
Reading the bibliography…
A line of recent works established that when training linear predictors over separable data, using gradient methods and exponentially-tailed losses, the predictors asymptotically converge in direction to the max-margin predictor.
A refined primal-dual analysis of the implicit bias
Ziwei Ji and Matus Telgarsky · 1906
Earlier work this paper cites.
Simplified pac-bayesian margin bounds
David McAllester · 2003
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
Characterizing implicit bias in terms of optimization geometry
Suriya Gunasekar · 2018
Cited alongside, same era.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason D Lee, Daniel Soudry, and Nati Srebro · 2018
Cited alongside, same era.
Lectures on convex optimization , volume 137
Yurii Nesterov · 2018
Cited alongside, same era.
Gradient descent aligns the layers of deep linear networks
Ziwei Ji and Matus Telgarsky
Cited in the paper.
Risk and parameter convergence of logistic regression
Ziwei Ji and Matus Telgarsky
Cited in the paper.
Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow relu networks
Ziwei Ji and Matus Telgarsky
Cited in the paper.
Convergence of gradient descent on separable data
Mor Shpigel Nacson, Jason Lee, Suriya Gunasekar, Pedro Henrique Pamplona Savarese, Nathan Srebro, and Daniel Soudry
Cited in the paper.
Stochastic gradient descent on separable data: Exact convergence with a fixed learning rate
Mor Shpigel Nacson, Nathan Srebro, and Daniel Soudry
Cited in the paper.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Later among the works it cites.
Gradient descent follows the regularization path for general losses
Miroslav Dudik, Ziwei Ji, Robert Schapire, and Matus Telgarsky · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…