Fetching the paper…
Reading the bibliography…
Gradient regularization (GR), which aims to penalize the gradient norm atop the loss function, has shown promising results in training modern over-parameterized deep neural networks.
Flat minima
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Probability models in engineering and science , volume 192
Benaroya, H., Han, S. M., and Nagurka, M · 2005
Earlier work this paper cites.
Neural networks for machine learning lecture 6a overview of mini-batch gradient descent
Hinton, G., Srivastava, N., and Swersky, K · 2012
Earlier work this paper cites.
Forecasting with moving averages
Nau, R · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G. E., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Ba, L. J., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
Training tips for the transformer model
Popel, M. and Bojar, O · 2018
Cited alongside, same era.
Implicit gradient regularization
Barrett, D. G. and Dherin, B · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
On the variance of the adaptive learning rate and beyond
Liu, L., Jiang, H., He, P., Chen, W., Liu, X., Gao, J., and Han, J · 2020
ASAM: adaptive sharpness-aware minimization for scale-invariant learning of deep neural networks
Kwon, J., Kim, J., Park, H., and Choi, I. K · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B · 2021
Later among the works it cites.
On the origin of implicit regularization in stochastic gradient descent
Smith, S. L., Dherin, B., Barrett, D. G., and De, S · 2021
Later among the works it cites.
Going deeper with image transformers
Touvron, H., Cord, M., Sablayrolles, A., Synnaeve, G., and Jégou, H · 2021
Later among the works it cites.
Penalizing gradient norm for efficiently improving generalization in deep learning
Zhao, Y., Zhang, H., and Hu, X · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Cited alongside, same era.
Efficient sharpness-aware minimization for improved training of neural networks
Du, J., Yan, H., Feng, J., Zhou, J. T., Zhen, L., Goh, R. S. M., and Tan, V. Y. F · 2021
Cited alongside, same era.
Sharpness-aware minimization for efficiently improving generalization
Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B · 2021
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I
Cited in the paper.
Randomized sharpness-aware training for boosting computational efficiency in deep learning
Zhao, Y., Zhang, H., and Hu, X
Cited in the paper.
Zhuang, J., Gong, B., Yuan, L., Cui, Y., Adam, H., Dvornek, N. C., Tatikonda, S., Duncan, J. S., and Liu, T · 2022
Later among the works it cites.
Understanding gradient regularization in deep learning: Efficient finite-difference computation and implicit bias
Karakida, R., Takase, T., Hayase, T., and Osawa, K · 2023
Later among the works it cites.
Samba: Regularized autoencoders perform sharpness-aware minimization
Reizinger, P. and Huszár, F · 2023
Later among the works it cites.