Fetching the paper…
Reading the bibliography…
Large learning rates, when applied to gradient descent for nonconvex optimization, yield various implicit biases including the edge of stability (Cohen et al., 2021), balancing (Wang et al., 2022), and catapult (Lewkowycz et al., 2020).
On richardson’s method for solving linear systems with positive definite matrices
D. Young · 1953
Earlier work this paper cites.
Growth conditions and the numerical range in a banach algebra
J. Stampfli and J. P. Williams · 1968
Earlier work this paper cites.
Growth conditions and regularity, a counterexample
M. Giaquinta · 1987
Earlier work this paper cites.
Quasiconvexity, growth conditions and partial regularity
M. Giaquinta · 1988
Earlier work this paper cites.
Keeping the neural networks simple by minimizing the description length of the weights
G. E. Hinton and D. Van Camp · 1993
Earlier work this paper cites.
Flat minima
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
(not) bounding the true error
J. Langford and R. Caruana · 2001
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Sharp minima can generalize for deep nets
L. Dinh, R. Pascanu, S. Bengio, and Y. Bengio · 2017
Earlier work this paper cites.
G. K. Dziugaite and D. M. Roy · 2017
Earlier work this paper cites.
Greed, hedging, and acceleration in convex optimization
J. Altschuler · 2018
Earlier work this paper cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
S. Arora, N. Cohen, and E. Hazan · 2018
Earlier work this paper cites.
Understanding batch normalization
N. Bjorck, C. P. Gomes, B. Selman, and K. Q. Weinberger · 2018
Earlier work this paper cites.
Error bounds, quadratic growth, and linear convergence of proximal methods
D. Drusvyatskiy and A. S. Lewis · 2018
Earlier work this paper cites.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
S. S. Du, W. Hu, and J. D. Lee · 2018
Earlier work this paper cites.
How does batch normalization help optimization?
S. Santurkar, D. Tsipras, A. Ilyas, and A. Madry · 2018
Earlier work this paper cites.
Towards flatter loss surface via nonmonotonic learning rate scheduling
S. Seong, Y. Lee, Y. Kee, D. Han, and J. Kim · 2018
Earlier work this paper cites.
L. N. Smith · 2018
Earlier work this paper cites.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
L. Wu, C. Ma, and W. E · 2018
Earlier work this paper cites.
An investigation into neural net optimization via hessian eigenvalue density
B. Ghorbani, S. Krishnan, and Y. Xiao · 2019
Cited alongside, same era.
The normalization method for alleviating pathological sharpness in wide neural networks
R. Karakida, S. Akaho, and S.-i. Amari · 2019
Cited alongside, same era.
Stochastic modified equations and dynamics of stochastic gradient algorithms i: Mathematical foundations
Q. Li, C. Tai, and E. Weinan · 2019
Cited alongside, same era.
Information-theoretic generalization bounds for sgld via data-dependent estimates
J. Negrea, M. Haghifam, G. K. Dziugaite, A. Khisti, and D. M. Roy · 2019
Cited alongside, same era.
Why gradient clipping accelerates training: A theoretical justification for adaptivity
J. Zhang, T. He, S. Sra, and A. Jadbabaie · 2019
Cited alongside, same era.
Understanding the generalization benefit of normalization layers: Sharpness reduction
K. Lyu, Z. Li, and S. Arora · 2022
Later among the works it cites.
Implicit bias of the step size in linear diagonal neural networks
M. S. Nacson, K. Ravichandran, N. Srebro, and D. Soudry · 2022
Later among the works it cites.
Learning threshold neurons via the “edge of stability”
K. Ahn, S. Bubeck, S. Chewi, Y. T. Lee, F. Suarez, and Y. Zhang · 2023
Closest in time.
Beyond the edge of stability via two-step gradient updates, 2023
L. Chen and J. Bruna · 2023
Closest in time.
Gradient descent finds the global optima of two-layer physics-informed neural networks
Y. Gao, Y. Gu, and M. Ng · 2023
Closest in time.
The inductive bias of flatness regularization for deep matrix factorization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Kong and M. Tao · 2020
Cited alongside, same era.
The large learning rate phase of deep learning: the catapult mechanism
A. Lewkowycz, Y. Bahri, E. Dyer, J. Sohl-Dickstein, and G. Gur-Ari · 2020
Cited alongside, same era.
Two-layer neural networks for partial differential equations: Optimization and generalization theory
T. Luo and H. Yang · 2020
Cited alongside, same era.
Gradient descent on neural networks typically occurs at the edge of stability
J. Cohen, S. Kaur, Y. Li, J. Z. Kolter, and A. Talwalkar · 2021
Cited alongside, same era.
Sharpness-aware minimization for efficiently improving generalization
P. Foret, A. Kleiner, H. Mobahi, and B. Neyshabur · 2021
Cited alongside, same era.
Catastrophic fisher explosion: Early phase fisher matrix impacts generalization
S. Jastrzebski, D. Arpit, O. Astrand, G. B. Kerg, H. Wang, C. Xiong, R. Socher, K. Cho, and K. J. Geras · 2021
Cited alongside, same era.
Beyond procrustes: Balancing-free gradient descent for asymmetric low-rank matrix sensing
C. Ma, Y. Li, and Y. Chi · 2021
Cited alongside, same era.
K. Gatmiry, Z. Li, C.-Y. Chuang, S. Reddi, T. Ma, and S. Jegelka · 2023
Closest in time.
Provably faster gradient descent via long steps
B. Grimmer · 2023
Closest in time.
Investigating the edge of stability phenomenon in reinforcement learning
R. Iordan, M. P. Deisenroth, and M. Rosca · 2023
Closest in time.
Phase diagram of early training dynamics in deep neural networks: effect of the learning rate, depth, and width
D. S. Kalra and M. Barkeshli · 2023
Closest in time.
I. Kreisler, M. S. Nacson, D. Soudry, and Y. Carmon · 2023
Closest in time.
Convex and non-convex optimization under generalized smoothness
H. Li, J. Qian, Y. Tian, A. Rakhlin, and A. Jadbabaie · 2023
Closest in time.
On progressive sharpening, flat minima and generalisation
L. E. MacDonald, J. Valmadre, and S. Lucey · 2023
Closest in time.
Trajectory alignment: Understanding the edge of stability phenomenon via bifurcation theory
M. Song and C. Yun · 2023
Closest in time.
Convergence of alternating gradient descent for matrix factorization
R. Ward and T. G. Kolda · 2023
Closest in time.
Sharpness minimization algorithms do not only minimize sharpness to achieve better generalization
K. Wen, T. Ma, and Z. Li · 2023
Closest in time.
Implicit bias of gradient descent for logistic regression at the edge of stability
J. Wu, V. Braverman, and J. D. Lee · 2023
Closest in time.
Y. Yang and D.-X. Zhou · 2023
Closest in time.
Salr: Sharpness-aware learning rate scheduler for improved generalization
X. Yue, M. Nouiehed, and R. Al Kontar · 2023
Closest in time.
Understanding edge-of-stability training dynamics with a minimalist example
X. Zhu, Z. Wang, X. Wang, M. Zhou, and R. Ge · 2023
Closest in time.