Fetching the paper…
Reading the bibliography…
We study the implicit bias of gradient descent methods in solving a binary classification problem over a linearly separable dataset.
Problem complexity and method efficiency in optimization
Arkadii Nemirovskii, David Borisovich Yudin, and Edgar Ronald Dawson · 1983
Earlier work this paper cites.
Efficient online and batch learning using forward backward splitting
John Duchi and Yoram Singer · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Earlier work this paper cites.
Stochastic convex optimization
Shai Shalev-Shwartz, Ohad Shamir, Nathan Srebro, and Karthik Sridharan · 2009
Earlier work this paper cites.
Convex analysis and nonlinear optimization: theory and examples
Jonathan Borwein and Adrian S Lewis · 2010
Earlier work this paper cites.
Dual averaging methods for regularized stochastic learning and online optimization
Lin Xiao · 2010
Earlier work this paper cites.
Non-strongly-convex smooth stochastic approximation with convergence rate O(1/n)
Francis Bach and Eric Moulines · 2013
Earlier work this paper cites.
Margins, shrinkage, and boosting
Matus Telgarsky · 2013
Cited alongside, same era.
Adaptivity of averaged stochastic gradient descent to local strong convexity for logistic regression
Francis R Bach · 2014
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2016
Cited alongside, same era.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
Alon Brutzkus, Amir Globerson, Eran Malach, and Shai Shalev-Shwartz · 2017
Cited alongside, same era.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Cited alongside, same era.
On the learning dynamics of deep neural networks
Remi Tachet des Combes, Mohammad Pezeshki, Samira Shabanian, Aaron Courville, and Yoshua Bengio · 2018
Closest in time.
Characterizing implicit bias in terms of optimization geometry
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro · 2018
Closest in time.
Risk and parameter convergence of logistic regression
Ziwei Ji and Matus Telgarsky · 2018
Closest in time.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Closest in time.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, and Nathan Srebro · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daniel Soudry, Elad Hoffer, and Nathan Srebro · 2017
Cited alongside, same era.
Convergence of gradient descent on separable data
Mor Shpigel Nacson, Jason Lee, Suriya Gunasekar, Nathan Srebro, and Daniel Soudry
Cited in the paper.
Stochastic gradient descent on separable data: Exact convergence with a fixed learning rate
Mor Shpigel Nacson, Nathan Srebro, and Daniel Soudry
Cited in the paper.
Closest in time.
Learning relu networks on linearly separable data: Algorithm, optimality, and generalization
Gang Wang, Georgios B Giannakis, and Jie Chen · 2018
Closest in time.