Fetching the paper…
Reading the bibliography…
The typical training of neural networks using large stepsize gradient descent (GD) under the logistic loss often involves two distinct phases, where the empirical risk oscillates in the first phase but decreases monotonically in the second phase.
On convergence proofs for perceptrons
Albert BJ Novikoff · 1962
Earlier work this paper cites.
Learning multiple layers of features from tiny images, 2009
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Accurate, large minibatch SGD: Training ImageNet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Earlier work this paper cites.
Train longer, generalize better: Closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Earlier work this paper cites.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
Alon Brutzkus, Amir Globerson, Eran Malach, and Shai Shalev-Shwartz · 2018
Earlier work this paper cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Earlier work this paper cites.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason D Lee, Daniel Soudry, and Nati Srebro · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Risk and parameter convergence of logistic regression
Ziwei Ji and Matus Telgarsky · 2018
Earlier work this paper cites.
Lectures on Convex Optimization
Yurii Nesterov · 2018
Earlier work this paper cites.
A mean field view of the landscape of two-layers neural networks
Mei Song, Andrea Montanari, and P Nguyen · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Cited alongside, same era.
How SGD selects the global minima in over-parameterized learning: A dynamical stability perspective
Lei Wu and Chao Ma · 2018
Cited alongside, same era.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lenaic Chizat and Francis Bach · 2020
Cited alongside, same era.
Gradient descent on neural networks typically occurs at the edge of stability
Jeremy Cohen, Simran Kaur, Yuanzhi Li, J Zico Kolter, and Ameet Talwalkar · 2020
Cited alongside, same era.
Directional convergence and alignment in deep learning
Ziwei Ji and Matus Telgarsky · 2020
Cited alongside, same era.
Stochasticity of deterministic gradient descent: Large learning rate for multiscale objective function
The asymmetric maximum margin bias of quasi-homogeneous neural networks
Daniel Kunin, Atsushi Yamamura, Chao Ma, and Surya Ganguli · 2022
Later among the works it cites.
Understanding the generalization benefit of normalization layers: Sharpness reduction
Kaifeng Lyu, Zhiyuan Li, and Sanjeev Arora · 2022
Later among the works it cites.
Beyond the quadratic approximation: The multiscale structure of neural network loss landscapes
Chao Ma, Daniel Kunin, Lei Wu, and Lexing Ying · 2022
Later among the works it cites.
Understanding edge-of-stability training dynamics with a minimalist example
Xingyu Zhu, Zixuan Wang, Xiang Wang, Mo Zhou, and Rong Ge · 2022
Later among the works it cites.
Learning threshold neurons via edge of stability
Kwangjun Ahn, Sebastien Bubeck, Sinho Chewi, Yin Tat Lee, Felipe Suarez, and Yi Zhang · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lingkai Kong and Molei Tao · 2020
Cited alongside, same era.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2020
Cited alongside, same era.
When does gradient descent with logistic loss interpolate using deep networks with smoothed relu activations?
Niladri S Chatterji, Philip M Long, and Peter Bartlett · 2021
Cited alongside, same era.
Provable generalization of sgd-trained neural networks of any width in the presence of adversarial label noise
Spencer Frei, Yuan Cao, and Quanquan Gu · 2021
Cited alongside, same era.
Fast margin maximization via dual acceleration
Ziwei Ji, Nathan Srebro, and Matus Telgarsky · 2021
Cited alongside, same era.
Understanding the unstable convergence of gradient descent
Kwangjun Ahn, Jingzhao Zhang, and Suvrit Sra · 2022
Cited alongside, same era.
Self-stabilization: The implicit bias of gradient descent at the edge of stability
Alex Damian, Eshaan Nichani, and Jason D Lee · 2022
Cited alongside, same era.
Beyond the edge of stability via two-step gradient updates
Lei Chen and Joan Bruna · 2023
Later among the works it cites.
Gradient descent monotonically decreases the sharpness of gradient flow solutions in scalar networks and beyond
Itai Kreisler, Mor Shpigel Nacson, Daniel Soudry, and Yair Carmon · 2023
Later among the works it cites.
Benign oscillation of stochastic gradient descent with large learning rate
Miao Lu, Beining Wu, Xiaodong Yang, and Difan Zou · 2023
Later among the works it cites.
Yuqing Wang, Zhenghao Xu, Tuo Zhao, and Molei Tao · 2023
Later among the works it cites.
Implicit bias of gradient descent for logistic regression at the edge of stability
Jingfeng Wu, Vladimir Braverman, and Jason D. Lee · 2023
Later among the works it cites.
Large stepsize gradient descent for logistic loss: Non-monotonicity of the loss improves optimization efficiency
Jingfeng Wu, Peter L Bartlett, Matus Telgarsky, and Bin Yu · 2024
Closest in time.