Fetching the paper…
Reading the bibliography…
Sharpness-aware minimization (SAM) was proposed to reduce sharpness of minima and has been shown to enhance generalization performance in various settings.
Simplifying neural nets by discovering flat minima
S. Hochreiter and J. Schmidhuber · 1994
Earlier work this paper cites.
Flat minima
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Intriguing properties of neural networks
C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus · 2014
Earlier work this paper cites.
Batch Normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition, 2015
K. Simonyan and A. Zisserman · 2015
Earlier work this paper cites.
Layer Normalization
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Wide residual networks
S. Zagoruyko and N. Komodakis · 2016
Earlier work this paper cites.
Sharp minima can generalize for deep nets
L. Dinh, R. Pascanu, S. Bengio, and Y. Bengio · 2017
Earlier work this paper cites.
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data
G. K. Dziugaite and D. M. Roy · 2017
Earlier work this paper cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
E. Hoffer, I. Hubara, and D. Soudry · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He · 2017
Earlier work this paper cites.
Understanding batch normalization
J. Bjorck, C. Gomes, B. Selman, and K. Q. Weinberger · 2018
Earlier work this paper cites.
The best of both worlds: Combining recent advances in neural machine translation
M. X. Chen, O. Firat, A. Bapna, M. Johnson, W. Macherey, G. Foster, L. Jones, M. Schuster, N. Shazeer, N. Parmar, A. Vaswani, J. Uszkoreit, L. Kaiser, Z. Chen, Y. Wu, and M. Hughes · 2018
Earlier work this paper cites.
How does batch normalization help optimization?
S. Santurkar, D. Tsipras, A. Ilyas, and A. Madry · 2018
Earlier work this paper cites.
Theoretical analysis of auto rate-tuning by Batch Normalization
S. Arora, Z. Li, and K. Lyu · 2019
Earlier work this paper cites.
Exponential convergence rates for Batch Normalization: The power of length-direction decoupling in non-convex optimization
J. Kohler, H. Daneshmand, A. Lucchi, M. Zhou, K. Neymeyr, and T. Hofmann · 2019
Earlier work this paper cites.
Revisit Batch Normalization: New understanding and refinement via composition optimization
X. Lian and J. Liu · 2019
Cited alongside, same era.
Decoupled weight decay regularization, 2019
I. Loshchilov and F. Hutter · 2019
Cited alongside, same era.
Pytorch image models
R. Wightman · 2019
Cited alongside, same era.
Understanding and improving layer normalization
J. Xu, X. Sun, Z. Zhang, G. Zhao, and J. Lin · 2019
Cited alongside, same era.
Fixup initialization: Residual learning without normalization
H. Zhang, Y. N. Dauphin, and T. Ma · 2019
Cited alongside, same era.
Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks
F. Croce and M. Hein · 2020
Cited alongside, same era.
Relative flatness and generalization
H. Petzka, M. Kamp, L. Adilova, C. Sminchisescu, and M. Boley · 2021
Later among the works it cites.
Sharp-MAML: Sharpness-aware model-agnostic meta learning
M. Abbas, Q. Xiao, L. Chen, P.-Y. Chen, and T. Chen · 2022
Later among the works it cites.
Towards understanding sharpness-aware minimization
M. Andriushchenko and N. Flammarion · 2022
Later among the works it cites.
Sharpness-aware minimization improves language model generalization
D. Bahri, H. Mobahi, and Y. Tay · 2022
Later among the works it cites.
When Vision Transformers outperform ResNets without pre-training or strong data augmentations
X. Chen, C.-J. Hsieh, and B. Gong · 2022
Later among the works it cites.
Efficient sharpness-aware minimization for improved training of neural networks
J. Du, H. Yan, J. Feng, J. T. Zhou, L. Zhen, R. S. M. Goh, and V. Tan · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A new look at ghost normalization
N. Dimitriou and O. Arandjelovic · 2020
Cited alongside, same era.
In search of robust measures of generalization
G. K. Dziugaite, A. Drouin, B. Neal, N. Rajkumar, E. Caballero, L. Wang, I. Mitliagkas, and D. M. Roy · 2020
Cited alongside, same era.
Fantastic generalization measures and where to find them
Y. Jiang*, B. Neyshabur*, H. Mobahi, D. Krishnan, and S. Bengio · 2020
Cited alongside, same era.
Four things everyone should know to improve batch normalization
C. Summers and M. J. Dinneen · 2020
Cited alongside, same era.
Intriguing properties of adversarial training at scale
C. Xie and A. Yuille · 2020
Cited alongside, same era.
PyHessian: Neural networks through the lens of the Hessian
Z. Yao, A. Gholami, K. Keutzer, and M. Mahoney · 2020
Cited alongside, same era.
Later among the works it cites.
When do flat minima optimizers work?
J. Kaddour, L. Liu, R. Silva, and M. J. Kusner · 2022
Later among the works it cites.
Fisher SAM: Information geometry and sharpness aware minimisation
M. Kim, D. Li, S. X. Hu, and T. Hospedales · 2022
Later among the works it cites.
Towards efficient and scalable sharpness-aware minimization
Y. Liu, S. Mai, X. Chen, C.-J. Hsieh, and Y. You · 2022
Later among the works it cites.
Understanding the generalization benefit of normalization layers: Sharpness reduction
K. Lyu, Z. Li, and S. Arora · 2022
Later among the works it cites.
Make sharpness-aware minimization stronger: A sparsified perturbation approach
P. Mi, L. Shen, T. Ren, Y. Zhou, X. Sun, R. Ji, and D. Tao · 2022
Later among the works it cites.
Train flat, then compress: Sharpness-aware minimization learns more compressible models
C. Na, S. V. Mehta, and E. Strubell · 2022
Later among the works it cites.
How to train your ViT? Data, augmentation, and regularization in Vision Transformers
A. Steiner, A. Kolesnikov, X. Zhai, R. Wightman, J. Uszkoreit, and L. Beyer · 2022
Later among the works it cites.
Revisiting learnable affines for Batch Norm in few-shot transfer learning
M. Yazdanpanah, A. A. Rahman, M. Chaudhary, C. Desrosiers, M. Havaei, E. Belilovsky, and S. E. Kahou · 2022
Later among the works it cites.
Surrogate gap minimization improves sharpness-aware training
J. Zhuang, B. Gong, L. Yuan, Y. Cui, H. Adam, N. Dvornek, S. Tatikonda, J. Duncan, and T. Liu · 2022
Later among the works it cites.
mSAM: Micro-batch-averaged sharpness-aware minimization
K. Behdin, Q. Song, A. Gupta, A. Acharya, D. Durfee, B. Ocejo, S. Keerthi, and R. Mazumder · 2023
Closest in time.
Symbolic discovery of optimization algorithms
X. Chen, C. Liang, D. Huang, E. Real, K. Wang, Y. Liu, H. Pham, X. Dong, T. Luong, C.-J. Hsieh, Y. Lu, and Q. V. Le · 2023
Closest in time.
SAM as an optimal relaxation of Bayes
T. Möllenhoff and M. E. Khan · 2023
Closest in time.
Sharpness-aware minimization alone can improve adversarial robustness
Z. Wei, J. Zhu, and Y. Zhang · 2023
Closest in time.