Fetching the paper…
Reading the bibliography…
Understanding the algorithmic bias of \emph{stochastic gradient descent} (SGD) is one of the key challenges in modern machine learning and deep learning theory.
Numerical linear algebra , volume 50
Lloyd N Trefethen and David Bau III · 1997
Earlier work this paper cites.
Stability and generalization
Olivier Bousquet and André Elisseeff · 2002
Earlier work this paper cites.
Stochastic approximation and recursive algorithms and applications , volume 35
Harold Kushner and G George Yin · 2003
Earlier work this paper cites.
Convex optimization
Stephen Boyd, Stephen P Boyd, and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Roman Vershynin · 2010
Earlier work this paper cites.
Matrix analysis
Roger A Horn and Charles R Johnson · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Ben Recht, and Yoram Singer · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Earlier work this paper cites.
On the diffusion approximation of nonconvex stochastic gradient descent
Wenqing Hu, Chris Junchi Li, Lei Li, and Jian-Guo Liu · 2017
Earlier work this paper cites.
Three factors influencing minima in sgd
Stanisław Jastrzębski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Cited alongside, same era.
Stochastic modified equations and adaptive stochastic gradient algorithms
Qianxiao Li, Cheng Tai, et al · 2017
Cited alongside, same era.
Stability and generalization of learning algorithms that converge to global optima
Zachary Charles and Dimitris Papailiopoulos · 2018
Cited alongside, same era.
Data-dependent stability of stochastic gradient descent
Ilja Kuzborskij and Christoph Lampert · 2018
Cited alongside, same era.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2018
Cited alongside, same era.
Gradient descent aligns the layers of deep linear networks
Ziwei Ji and Matus Jan Telgarsky · 2019
Later among the works it cites.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Yuanzhi Li, Colin Wei, and Tengyu Ma · 2019
Later among the works it cites.
On the heavy-tailed theory of stochastic gradient descent for deep neural networks
Umut Simsekli, Mert Gürbüzbalaban, Thanh Huy Nguyen, Gaël Richard, and Levent Sagun · 2019
Later among the works it cites.
The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma · 2019
Later among the works it cites.
The implicit regularization of stochastic gradient flow for least squares
Alnur Ali, Edgar Dobriban, and Ryan J Tibshirani · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Implicit regularization in nonconvex statistical estimation: Gradient descent converges linearly for phase retrieval and matrix completion
Cong Ma, Kaizheng Wang, Yuejie Chi, and Yuxin Chen · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Cited alongside, same era.
Connecting optimization and regularization paths
Arun Suggala, Adarsh Prasad, and Pradeep K Ravikumar · 2018
Cited alongside, same era.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
Lei Wu, Chao Ma, and E Weinan · 2018
Cited alongside, same era.
Global convergence of langevin dynamics based algorithms for nonconvex optimization
Pan Xu, Jinghui Chen, Difan Zou, and Quanquan Gu · 2018
Cited alongside, same era.
A continuous-time view of early stopping for least squares regression
Alnur Ali, J Zico Kolter, and Ryan J Tibshirani · 2019
Cited alongside, same era.
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 2019
Cited alongside, same era.
Closest in time.
Benign overfitting in linear regression
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler · 2020
Closest in time.
Stability of stochastic gradient descent on nonsmooth convex losses
Raef Bassily, Vitaly Feldman, Cristóbal Guzmán, and Kunal Talwar · 2020
Closest in time.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lenaic Chizat and Francis Bach · 2020
Closest in time.
Shape matters: Understanding the implicit bias of the noise covariance
Jeff Z HaoChen, Colin Wei, Jason D Lee, and Tengyu Ma · 2020
Closest in time.
Gradient descent follows the regularization path for general losses
Ziwei Ji, Miroslav Dudík, Robert E Schapire, and Matus Telgarsky · 2020
Closest in time.
The two regimes of deep network training
Guillaume Leclerc and Aleksander Madry · 2020
Closest in time.
The large learning rate phase of deep learning: the catapult mechanism
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2020
Closest in time.
Implicit bias in deep linear classification: Initialization scale vs training accuracy
Edward Moroshko, Suriya Gunasekar, Blake Woodworth, Jason D Lee, Nathan Srebro, and Daniel Soudry · 2020
Closest in time.
Learning rate annealing can provably help generalization, even for convex problems
Preetum Nakkiran · 2020
Closest in time.
On the noisy gradient descent that generalizes as sgd
Jingfeng Wu, Wenqing Hu, Haoyi Xiong, Jun Huan, Vladimir Braverman, and Zhanxing Zhu · 2020
Closest in time.