Fetching the paper…
Reading the bibliography…
This paper shows that the implicit bias of gradient descent on linearly separable data is exactly characterized by the optimal solution of a dual optimization problem given by a smoothed margin, even for general losses.
Inequalities
G. H. Hardy, J. E. Littlewood, and G. Pólya · 1934
Earlier work this paper cites.
Convex analysis
R Tyrrell Rockafellar · 1970
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E. Schapire · 1997
Earlier work this paper cites.
Boosting the margin: A new explanation for the effectiveness of voting methods
Robert E. Schapire, Yoav Freund, Peter Bartlett, and Wee Sun Lee · 1997
Earlier work this paper cites.
Approximate solutions to markov decision processes
Geoffrey J Gordon · 1999
Earlier work this paper cites.
Boosting as entropy projection
Jyrki Kivinen and Manfred K. Warmuth · 1999
Earlier work this paper cites.
Convex Analysis and Nonlinear Optimization
Jonathan Borwein and Adrian Lewis · 2000
Earlier work this paper cites.
Fundamentals of Convex Analysis
Jean-Baptiste Hiriart-Urruty and Claude Lemaréchal · 2001
Earlier work this paper cites.
Logistic regression, AdaBoost and Bregman distances
Michael Collins, Robert E. Schapire, and Yoram Singer · 2002
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course
Yurii Nesterov · 2004
Earlier work this paper cites.
The dynamics of adaboost: Cyclic behavior and convergence of margins
Cynthia Rudin, Ingrid Daubechies, and Robert E Schapire · 2004
Cited alongside, same era.
Online learning: Theory, algorithms, and applications
Shai Shalev-Shwartz and Yoram Singer · 2007
Cited alongside, same era.
On the equivalence of weak learnability and linear separability: New relaxations and efficient boosting algorithms
Shai Shalev-Shwartz and Yoram Singer · 2008
Cited alongside, same era.
The convergence rate of AdaBoost
Robert E. Schapire · 2010
Cited alongside, same era.
Online learning and online convex optimization
Shai Shalev-Shwartz et al · 2011
Cited alongside, same era.
Sublinear optimization for machine learning
Kenneth L Clarkson, Elad Hazan, and David P Woodruff · 2012
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Later among the works it cites.
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Later among the works it cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2017
Later among the works it cites.
Convergence of gradient descent on separable data
Mor Shpigel Nacson, Jason Lee, Suriya Gunasekar, Nathan Srebro, and Daniel Soudry · 2018
Later among the works it cites.
Gradient descent maximizes the margin of homogeneous neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Boosting: Foundations and Algorithms
Robert E. Schapire and Yoav Freund · 2012
Cited alongside, same era.
Adaboost and forward stagewise regression are first-order convex optimization methods
Robert M Freund, Paul Grigas, and Rahul Mazumder · 2013
Cited alongside, same era.
Margins, shrinkage, and boosting
Matus Telgarsky · 2013
Cited alongside, same era.
Characterizing implicit bias in terms of optimization geometry
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro
Cited in the paper.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason D Lee, Daniel Soudry, and Nati Srebro
Cited in the paper.
Gradient descent aligns the layers of deep linear networks
Ziwei Ji and Matus Telgarsky
Cited in the paper.
Kaifeng Lyu and Jian Li · 2019
Closest in time.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lenaic Chizat and Francis Bach · 2020
Closest in time.
Directional convergence and alignment in deep learning
Ziwei Ji and Matus Telgarsky · 2020
Closest in time.
Gradient descent follows the regularization path for general losses
Ziwei Ji, Miroslav Dudík, Robert E Schapire, and Matus Telgarsky · 2020
Closest in time.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Closest in time.