Fetching the paper…
Reading the bibliography…
It is well known that the finite step-size ($h$) in Gradient Descent (GD) implicitly regularizes solutions to flatter minima.
Some methods of speeding up the convergence of iteration methods
Boris T Polyak · 1964
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
Geoffrey E. Hinton, Simon Osindero, and Yee Whye Teh · 2006
Earlier work this paper cites.
Scaling learning algorithms towards AI
Yoshua Bengio and Yann LeCun · 2007
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman, Geoffrey Hinton, et al · 2012
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Densenet: Implementing efficient convnet descriptor pyramids
Forrest Iandola, Matt Moskewicz, Sergey Karayev, Ross Girshick, Trevor Darrell, and Kurt Keutzer · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
Deep learning , volume 1
Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Cited alongside, same era.
Stochastic modified equations and adaptive stochastic gradient algorithms
Qianxiao Li, Cheng Tai, and E Weinan · 2017
Cited alongside, same era.
Implicit regularization in deep learning
Behnam Neyshabur · 2017
Cited alongside, same era.
A bayesian perspective on generalization and stochastic gradient descent
Samuel L Smith and Quoc V Le · 2017
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Cited alongside, same era.
Theoretical issues in deep networks
Tomaso Poggio, Andrzej Banburski, and Qianli Liao · 2020
Later among the works it cites.
Implicit regularization in deep learning may not be explainable by norms
Noam Razin and Nadav Cohen · 2020
Later among the works it cites.
On the noisy gradient descent that generalizes as sgd
Jingfeng Wu, Wenqing Hu, Haoyi Xiong, Jun Huan, Vladimir Braverman, and Zhanxing Zhu · 2020
Later among the works it cites.
Gradient descent on neural networks typically occurs at the edge of stability
Jeremy M Cohen, Simran Kaur, Yuanzhi Li, J Zico Kolter, and Ameet Talwalkar · 2021
Later among the works it cites.
Continuous time analysis of momentum methods
Nikola B Kovachki and Andrew M Stuart · 2021
Later among the works it cites.
On the origin of implicit regularization in stochastic gradient descent
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 2019
Cited alongside, same era.
The implicit bias of gradient descent on nonseparable data
Ziwei Ji and Matus Telgarsky · 2019
Cited alongside, same era.
Stochastic modified equations and dynamics of stochastic gradient algorithms i: Mathematical foundations
Qianxiao Li, Cheng Tai, and Weinan E · 2019
Cited alongside, same era.
Implicit gradient regularization
David GT Barrett and Benoit Dherin · 2020
Cited alongside, same era.
Explicit regularisation in gaussian noise injections
Alexander Camuto, Matthew Willetts, Umut Simsekli, Stephen J Roberts, and Chris C Holmes · 2020
Cited alongside, same era.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2020
Cited alongside, same era.
Anticorrelated noise injection for improved generalization
Antonio Orvieto, Hans Kersting, Frank Proske, Francis Bach, and Aurelien Lucchi
Cited in the paper.
Samuel L Smith, Benoit Dherin, David GT Barrett, and Soham De · 2021
Later among the works it cites.
Implicit regularization in relu networks with the square loss
Gal Vardi and Ohad Shamir · 2021
Later among the works it cites.
Does momentum change the implicit regularization on separable data?, 2021
Bohan Wang, Qi Meng, Huishuai Zhang, Ruoyu Sun, Wei Chen, Zhi-Ming Ma, and Tie-Yan Liu · 2021
Later among the works it cites.
Quasi-potential theory for escape problem: Quantitative sharpness effect on SGD’s escape from local minima, 2022
Hikaru Ibayashi and Masaaki Imaizumi · 2022
Later among the works it cites.
Towards understanding how momentum improves generalization in deep learning
Samy Jelassi and Yuanzhi Li · 2022
Later among the works it cites.