Fetching the paper…
Reading the bibliography…
We examine gradient descent on unregularized logistic regression problems, with homogeneous linear predictors on linearly separable datasets.
Boosting the margin: A new explanation for the effectiveness of voting methods
Robert E Schapire, Yoav Freund, Peter Bartlett, Wee Sun Lee, et al · 1998
Earlier work this paper cites.
Margin Maximizing Loss Functions
Saharon Rosset, Ji Zhu, and Trevor J Hastie · 2004
Earlier work this paper cites.
Boosting with early stopping: Convergence and consistency
Tong Zhang, Bin Yu, et al · 2005
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Margins, shrinkage and boosting
Matus Telgarsky · 2013
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
EE6151, Convex optimization algorithms. Unconstrained minimization: Gradient descent algorithm, 2015
Radha Krishna Ganti · 2015
Earlier work this paper cites.
Adam: a Method for Stochastic Optimization
Diederik P Kingma and Jimmy Lei Ba · 2015
Earlier work this paper cites.
Path-sgd: Path-normalized optimization in deep neural networks
Behnam Neyshabur, Ruslan R Salakhutdinov, and Nati Srebro · 2015
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Benjamin Recht, and Y Singer · 2016
Cited alongside, same era.
Implicit Regularization in Matrix Factorization
Suriya Gunasekar, Blake Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nathan Srebro · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and D. Soudry · 2017
Cited alongside, same era.
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Cited alongside, same era.
Exploring Generalization in Deep Learning
Implicit Bias of Gradient Descent on Linear Convolutional Networks
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro · 2018
Closest in time.
Characterizing implicit bias in terms of optimization geometry
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro · 2018
Closest in time.
Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
I Hubara, M Courbariaux, D. Soudry, R El-yaniv, and Y Bengio · 2018
Closest in time.
Risk and parameter convergence of logistic regression
Ziwei Ji and Matus Telgarsky · 2018
Closest in time.
The Implicit Bias of Gradient Descent on Separable Data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, and N Srebro · 2018
Closest in time.
Stochastic Gradient Descent on Separable Data Exact Convergence with a Fixed Learning Rate
Mor Shpigel Nacson, Nati Srebro, and Daniel Soudry · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nathan Srebro · 2017
Cited alongside, same era.
The Marginal Value of Adaptive Gradient Methods in Machine Learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nathan Srebro, and Benjamin Recht · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Cited alongside, same era.
Closest in time.
Convergence of Gradient Descent on Separable Data
Mor Shpigel Nacson, Jason Lee, Suriya Gunasekar, Nathan Srebro, and Daniel Soudry · 2019
Closest in time.
The implicit bias of gradient descent on separable multiclass data
Hrithik Ravi, Clayton Scott, Daniel Soudry, and Yutong Wang · 2024
Closest in time.