Fetching the paper…
Reading the bibliography…
A quadratic approximation of neural network loss landscapes has been extensively used to study the optimization process of these networks.
Second order properties of error surfaces: Learning time and generalization
Yann LeCun, Ido Kanter, and Sara Solla · 1990
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Earlier work this paper cites.
Towards theoretical understanding of large batch training in stochastic gradient descent
Xiaowu Dai and Yuhua Zhu · 2018
Earlier work this paper cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Earlier work this paper cites.
Neural tangent kernel: convergence and generalization in neural networks (invited paper)
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
An alternative view: When does sgd escape local minima?
Bobby Kleinberg, Yuanzhi Li, and Yang Yuan · 2018
Earlier work this paper cites.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
Lei Wu, Chao Ma, and Weinan E · 2018
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
A comparative analysis of the optimization and generalization property of two-layer neural network and random feature models under gradient descent dynamics
Weinan E, Chao Ma, and Lei Wu · 2019
Cited alongside, same era.
Niv Giladi, Mor Shpigel Nacson, Elad Hoffer, and Daniel Soudry · 2019
Cited alongside, same era.
Loss landscape sightseeing with multi-point optimization
Ivan Skorokhodov and Mikhail Burtsev · 2019
Cited alongside, same era.
Gradient descent on neural networks typically occurs at the edge of stability
Jeremy M Cohen, Simran Kaur, Yuanzhi Li, J Zico Kolter, and Ameet Talwalkar · 2021
Later among the works it cites.
Label noise sgd provably prefers flat global minimizers
Alex Damian, Tengyu Ma, and Jason D Lee · 2021
Later among the works it cites.
Limiting dynamics of sgd: Modified loss, phase space oscillations, and anomalous diffusion
Daniel Kunin, Javier Sagastuy-Brena, Lauren Gillespie, Eshed Margalit, Hidenori Tanaka, Surya Ganguli, and Daniel LK Yamins · 2021
Later among the works it cites.
What happens after sgd reaches zero loss?–a mathematical framework
Zhiyuan Li, Tianhao Wang, and Sanjeev Arora · 2021
Later among the works it cites.
On linear stability of sgd and input-smoothness of neural networks
Chao Ma and Lexing Ying · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kaichao You, Mingsheng Long, Jianmin Wang, and Michael I Jordan · 2019
Cited alongside, same era.
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
Guy Blanc, Neha Gupta, Gregory Valiant, and Paul Valiant · 2020
Cited alongside, same era.
Stochasticity of deterministic gradient descent: Large learning rate for multiscale objective function
Lingkai Kong and Molei Tao · 2020
Cited alongside, same era.
Gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2020
Cited alongside, same era.
Later among the works it cites.
Large learning rate tames homogeneity: Convergence and balancing effect
Yuqing Wang, Minshuo Chen, Tuo Zhao, and Molei Tao · 2021
Later among the works it cites.
Understanding the unstable convergence of gradient descent
Kwangjun Ahn, Jingzhao Zhang, and Suvrit Sra · 2022
Closest in time.