Fetching the paper…
Reading the bibliography…
We present a new approach to understanding the relationship between loss curvature and input-output model behaviour in deep learning.
Improving generalization performance using double backpropagation
Harris Drucker and Yann Le Cun · 1992
Earlier work this paper cites.
Flat minima
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Stability and Generalization
O. Bousquet and A. Elisseeff · 2002
Earlier work this paper cites.
Concise formulas for the area and volume of a hyperspherical cap
S. Li · 2011
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
M. Hardt, B. Recht, and Y. Singer · 2016
Earlier work this paper cites.
The covering radius of randomly distributed points on a manifold
A. Reznikov and E. B. Saff · 2016
Earlier work this paper cites.
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, P. Tak, and P. Tang · 2017
Earlier work this paper cites.
Entropy-SGD: Biasing Gradient Descent into Wide Valleys
P. Chaudhari, A. Choromanska, S. Soatto, Y. LeCun, C. Baldassi, C. Borgs, J. Chayes, L. Sagun, and R. Zecchina · 2017
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
P. Bartlett, D. J. Foster, and M. Telgarsky · 2017
Earlier work this paper cites.
Exploring Generalization in Deep Learning
B. Neyshabur, S. Bhojanapalli, D. Mcallester, and N. Srebro · 2017
Earlier work this paper cites.
Sharp minima can generalize for deep nets
L. Dinh, R. Pascanu, S. Bengio, and Y. Bengio · 2017
Earlier work this paper cites.
Computing Nonvacuous Generalization Bounds for Deep (Stochastic) Neural Networks with Many More Parameters than Training Data
G. K. Dziugaite and D. M. Roy · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Earlier work this paper cites.
Sensitivity and generalization in neural networks: an empirical study
Roman Novak, Yasaman Bahri, Daniel A Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Earlier work this paper cites.
The Full Spectrum of Deepnet Hessians at Scale: Dynamics with SGD Training and Sample Size
V. Papyan · 2018
Earlier work this paper cites.
Neural Tangent Kernel: Convergence and Generalization in Neural Networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Earlier work this paper cites.
How SGD Selects the Global Minima in Over-parameterized Learning: A Dynamical Stability Perspective
L. Wu, C. Ma, and W. E · 2018
Earlier work this paper cites.
Stability and Generalization of Learning Algorithms that Converge to Global Optima
Z. Charles and D. Papailiopoulos · 2018
Earlier work this paper cites.
Data-Dependent Stability of Stochastic Gradient Descent
I. Kuzborskij and C. H. Lampert · 2018
Cited alongside, same era.
Spectral Normalization for Generative Adversarial Networks
T. Miyato, T. Kataoka, M. Koyama, and Y. Yoshida · 2018
Cited alongside, same era.
mixup: Beyond Empirical Risk Minimization
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz · 2018
Cited alongside, same era.
Robust Learning with Jacobian Regularization
J. Hoffman, D. A. Roberts, and S. Yaida · 2019
Cited alongside, same era.
Measurements of three-level hierarchical structure in the outliers in the spectrum of deepnet hessians
V. Papyan · 2019
Cited alongside, same era.
An Investigation into Neural Net Optimization via Hessian Eigenvalue Density
A Universal Law of Robustness via Isoperimetry
S. Bubeck and M. Sellke · 2021
Later among the works it cites.
Regularisation of neural networks by enforcing Lipschitz continuity
H. Gouk, E. Frank, B. Pfahringer, and M. J. Cree · 2021
Later among the works it cites.
On linear stability of sgd and input-smoothness of neural networks
C. Ma and L. Ying · 2021
Later among the works it cites.
The Implicit Bias of Minima Stability: A View from Function Space
R. Mulayoff, T. Michaeli, and D. Soudry · 2021
Later among the works it cites.
Relative flatness and generalization
Henning Petzka, Michael Kamp, Linara Adilova, Cristian Sminchisescu, and Mario Boley · 2021
Later among the works it cites.
Understanding the Generalization Benefit of Normalization Layers: Sharpness Reduction
K. Lyu, Z. Li, and S. Arora · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Ghorbani, S. Krishnan, and Y. Xiao · 2019
Cited alongside, same era.
The Anisotropic Noise in Stochastic Gradient Descent: Its Behavior of Escaping from Sharp Minima and Regularization Effects
Z. Zhu, J. Wu, B. Yu, L. Wu, and J. Ma · 2019
Cited alongside, same era.
Control Batch Size and Learning Rate to Generalize Well: Theoretical and Empirical Evidence
F. He and T. Liu and D. Tao · 2019
Cited alongside, same era.
Data-dependent sample complexity of deep neural networks via lipschitz augmentation
C. Wei and T. Ma · 2019
Cited alongside, same era.
A Convergence Theory for Deep Learning via Over-Parameterization
Z. Allen-Zhu, Y. Li, and Z. Song · 2019
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
J. Lee, L. Xiao, S. Schoenholtz, Y. Bahri, R. Novak, J. Sohl-Dickstein, and J. Pennington · 2019
Cited alongside, same era.
Efficient and accurate estimation of lipschitz constants for deep neural networks
M. Fazlyab, A. Robey, H. Hassani, M. Morari, and G. Pappas · 2019
Cited alongside, same era.
A reparametrization-invariant sharpness measure based on information geometry
C. Jang, S. Lee, F. Park, and Y.-K. Noh · 2022
Later among the works it cites.
A global analysis of global optimisation
L. E. MacDonald, H. Saratchandran, J. Valmadre, and S. Lucey · 2022
Later among the works it cites.
Understanding Gradient Descent on Edge of Stability in Deep Learning
S. Arora, Z. Li, and A. Panigrahi · 2022
Later among the works it cites.
Analyzing Sharpness along GD Trajectory: Progressive Sharpening and Edge of Stability
Z. Wang, Z. Li, and J. Li · 2022
Later among the works it cites.
Understanding the unstable convergence of gradient descent
K. Ahn, J. Zhang, and S. Sra · 2022
Later among the works it cites.
How much does Initialization Affect Generalization?
S. Ramasinghe, L. E. MacDonald, M. Farazi, H. Saratchandran, and S. Lucey · 2023
Closest in time.
On the lipschitz constant of deep networks and double descent
M. Gamba, H. Azizpour, and M. Bjorkman · 2023
Closest in time.
Implicit jacobian regularization weighted with impurity of probability output
S. Lee, J. Park, and J. Lee · 2023
Closest in time.
Sparsity-aware generalization theory for deep neural networks
R. Muthukumar and J. Sulam · 2023
Closest in time.
A new characterization of the edge of stability based on a sharpness measure aware of batch gradient distribution
S. Lee and C. Jang · 2023
Closest in time.
Understanding Edge-of-Stability Training Dynamics with a Minimalist Example
X. Zhu, Z. Wang, X. Wang, M. Zhou, and R. Ge · 2023
Closest in time.
Self-Stabilization: The Implicit Bias of Gradient Descent at the Edge of Stability
A. Damian, E. Nichani, and J. Lee · 2023
Closest in time.