Fetching the paper…
Reading the bibliography…
Gradient Descent (GD) is a powerful workhorse of modern machine learning thanks to its scalability and efficiency in high-dimensional spaces.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Introductory lectures on convex programming, 1998
Yu Nesterov · 1998
Earlier work this paper cites.
Efficient backprop
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Understanding batch normalization
Nils Bjorck, Carla P Gomes, Bart Selman, and Kilian Q Weinberger · 2018
Earlier work this paper cites.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Earlier work this paper cites.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
Simon S Du, Wei Hu, and Jason D Lee · 2018
Earlier work this paper cites.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2018
Earlier work this paper cites.
Fantastic generalization measures and where to find them
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 2019
Earlier work this paper cites.
Implicit gradient regularization
David Barrett and Benoit Dherin · 2020
Cited alongside, same era.
Gradient descent on neural networks typically occurs at the edge of stability
Jeremy Cohen, Simran Kaur, Yuanzhi Li, J Zico Kolter, and Ameet Talwalkar · 2020
Cited alongside, same era.
The large learning rate phase of deep learning: the catapult mechanism
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2020
Cited alongside, same era.
Learning a single neuron with gradient methods
Gilad Yehudai and Shamir Ohad · 2020
Cited alongside, same era.
Label noise sgd provably prefers flat global minimizers
Alex Damian, Tengyu Ma, and Jason D Lee · 2021
Cited alongside, same era.
Implicit regularization in relu networks with the square loss
Gal Vardi and Ohad Shamir · 2021
Later among the works it cites.
Learning a single neuron with bias using gradient descent
Gal Vardi, Gilad Yehudai, and Ohad Shamir · 2021
Later among the works it cites.
Large learning rate tames homogeneity: Convergence and balancing effect
Yuqing Wang, Minshuo Chen, Tuo Zhao, and Molei Tao · 2021
Later among the works it cites.
Global convergence of gradient descent for asymmetric low-rank matrix factorization
Tian Ye and Simon S Du · 2021
Later among the works it cites.
Understanding the unstable convergence of gradient descent
Kwangjun Ahn, Jingzhao Zhang, and Suvrit Sra · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Omer Elkabetz and Nadav Cohen · 2021
Cited alongside, same era.
A loss curvature perspective on training instability in deep learning
Justin Gilmer, Behrooz Ghorbani, Ankush Garg, Sneha Kudugunta, Behnam Neyshabur, David Cardoze, George Dahl, Zachary Nado, and Orhan Firat · 2021
Cited alongside, same era.
Catastrophic fisher explosion: Early phase fisher matrix impacts generalization
Stanislaw Jastrzebski, Devansh Arpit, Oliver Astrand, Giancarlo B Kerg, Huan Wang, Caiming Xiong, Richard Socher, Kyunghyun Cho, and Krzysztof J Geras · 2021
Cited alongside, same era.
On nonconvex optimization for machine learning: Gradients, stochasticity, and saddle points
Chi Jin, Praneeth Netrapalli, Rong Ge, Sham M Kakade, and Michael I Jordan · 2021
Cited alongside, same era.
The sobolev regularization effect of stochastic gradient descent
Chao Ma and Lexing Ying · 2021
Cited alongside, same era.
On the origin of implicit regularization in stochastic gradient descent
Samuel L Smith, Benoit Dherin, David GT Barrett, and Soham De · 2021
Cited alongside, same era.
Sanjeev Arora, Zhiyuan Li, and Abhishek Panigrahi · 2022
Closest in time.
Self-stabilization: The implicit bias of gradient descent at the edge of stability
Alex Damian, Eshaan Nichani, and Jason D Lee · 2022
Closest in time.
Flat minima generalize for low-rank matrix recovery
Lijun Ding, Dmitriy Drusvyatskiy, and Maryam Fazel · 2022
Closest in time.
Understanding the generalization benefit of normalization layers: Sharpness reduction
Kaifeng Lyu, Zhiyuan Li, and Sanjeev Arora · 2022
Closest in time.
The multiscale structure of neural network loss functions: The effect on optimization and origin
Chao Ma, Lei Wu, and Lexing Ying · 2022
Closest in time.