Fetching the paper…
Reading the bibliography…
Sharpness of minima is a promising quantity that can correlate with generalization in deep networks and, when optimized during training, can improve generalization.
Simplifying neural nets by discovering flat minima
Hochreiter, S. and Schmidhuber, J · 1995
Earlier work this paper cites.
Efficient backprop
LeCun, Y. A., Bottou, L., Orr, G. B., and Müller, K.-R · 2012
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J., Sagun, L., and Zecchina, R · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y · 2017
Earlier work this paper cites.
Three factors influencing minima in sgd
Jastrzebski, S., Kenton, Z., Arpit, D., Ballas, N., Fischer, A., Bengio, Y., and Storkey, A · 2017
Earlier work this paper cites.
Exploring generalization in deep learning
Neyshabur, B., Bhojanapalli, S., McAllester, D., and Srebro, N · 2017
Earlier work this paper cites.
Entropy-sgd optimizes the prior of a pac-bayes bound: Generalization properties of entropy-sgd and data-dependent priors
Dziugaite, G. K. and Roy, D · 2018
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D., and Wilson, A. G · 2018
Earlier work this paper cites.
A Bayesian perspective on generalization and stochastic gradient descent
Smith, S. L. and Le, Q. V · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A., Nangia, N., and Bowman, S · 2018
Earlier work this paper cites.
Xing, C., Arpit, D., Tsirigotis, C., and Bengio, Y · 2018
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D · 2018
Earlier work this paper cites.
Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models
Barbu, A., Mayo, D., Alverio, J., Luo, W., Wang, C., Gutfreund, D., Tenenbaum, J., and Katz, B · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Benchmarking neural network robustness to common corruptions and perturbations
Hendrycks, D. and Dietterich, T · 2019
Earlier work this paper cites.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Li, Y., Wei, C., and Ma, T · 2019
Earlier work this paper cites.
Fisher-rao metric, geometry, and complexity of neural networks
Liang, T., Poggio, T., Rakhlin, A., and Stokes, J · 2019
Earlier work this paper cites.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
McCoy, T., Pavlick, E., and Linzen, T · 2019
Earlier work this paper cites.
Do imagenet classifiers generalize to imagenet?
Recht, B., Roelofs, R., Schmidt, L., and Shankar, V · 2019
Earlier work this paper cites.
Learning robust global representations by penalizing local predictive power
Wang, H., Ge, S., Lipton, Z., and Xing, E. P · 2019
Earlier work this paper cites.
Implicit regularization for deep neural networks driven by an Ornstein-Uhlenbeck like process
Blanc, G., Gupta, N., Valiant, G., and Valiant, P · 2020
Cited alongside, same era.
Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks
Croce, F. and Hein, M · 2020
Cited alongside, same era.
Randaugment: Practical automated data augmentation with a reduced search space
Cubuk, E. D., Zoph, B., Shlens, J., and Le, Q. V · 2020
Cited alongside, same era.
In search of robust measures of generalization
Dziugaite, G. K., Drouin, A., Neal, B., Rajkumar, N., Caballero, E., Wang, L., Mitliagkas, I., and Roy, D. M · 2020
Cited alongside, same era.
Granziol, D · 2020
Cited alongside, same era.
Relative flatness and generalization
Petzka, H., Kamp, M., Adilova, L., Sminchisescu, C., and Boley, M · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Later among the works it cites.
On the origin of implicit regularization in stochastic gradient descent
Smith, S. L., Dherin, B., Barrett, D. G., and De, S · 2021
Later among the works it cites.
How to train your vit? data, augmentation, and regularization in vision transformers
Steiner, A., Kolesnikov, A., Zhai, X., Wightman, R., Uszkoreit, J., and Beyer, L · 2021
Later among the works it cites.
Relating adversarially robust generalization to flat minima
Stutz, D., Hein, M., and Schiele, B · 2021
Later among the works it cites.
An empirical investigation of domain generalization with empirical risk minimizers
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jiang, Y., Neyshabur, B., Mobahi, H., Krishnan, D., and Bengio, S · 2020
Cited alongside, same era.
Reconciling modern deep learning with traditional optimization analyses: The intrinsic learning rate
Li, Z., Lyu, K., and Arora, S · 2020
Cited alongside, same era.
BERTs of a feather do not generalize together: Large variability in generalization across models with similar test set performance
McCoy, R. T., Min, J., and Linzen, T · 2020
Cited alongside, same era.
Normalized flat minima: Exploring scale invariant definition of flat minima for neural networks using pac-bayesian analysis
Tsuzuku, Y., Sato, I., and Sugiyama, M · 2020
Cited alongside, same era.
Kernel and rich regimes in overparametrized models
Woodworth, B., Gunasekar, S., Lee, J. D., Moroshko, E., Savarese, P., Golan, I., Soudry, D., and Srebro, N · 2020
Cited alongside, same era.
Adversarial weight perturbation helps robust generalization
Wu, D., Xia, S.-t., and Wang, Y · 2020
Cited alongside, same era.
Towards theoretically understanding why SGD generalizes better than Adam in deep learning
Zhou, P., Feng, J., Ma, C., Xiong, C., Hoi, S. C. H., et al · 2020
Cited alongside, same era.
Vedantam, S. R., Lopez-Paz, D., and Schwab, D. J · 2021
Later among the works it cites.
Why flatness does and does not correlate with generalization for deep neural networks
Zhang, S., Reid, I., Pérez, G. V., and Louis, A · 2021
Later among the works it cites.
Regularizing neural networks via adversarial model perturbation
Zheng, Y., Zhang, R., and Mao, Y · 2021
Later among the works it cites.
Towards understanding sharpness-aware minimization
Andriushchenko, M. and Flammarion, N · 2022
Later among the works it cites.
Understanding gradient descent on edge of stability in deep learning
Arora, S., Li, Z., and Panigrahi, A · 2022
Later among the works it cites.
Low-pass filtering sgd for recovering flat optima in the deep learning optimization landscape
Bisla, D., Wang, J., and Choromanska, A · 2022
Later among the works it cites.
When vision transformers outperform resnets without pre-training or strong data augmentations?
Chen, X., Hsieh, C.-J., and Gong, B · 2022
Later among the works it cites.
Sharpness-aware training for free
Du, J., Daquan, Z., Feng, J., Tan, V., and Zhou, J. T · 2022
Later among the works it cites.
On the maximum hessian eigenvalue and generalization
Kaur, S., Cohen, J., and Lipton, Z. C · 2022
Later among the works it cites.
Understanding the generalization benefit of normalization layers: Sharpness reduction
Lyu, K., Li, Z., and Arora, S · 2022
Later among the works it cites.
How do vision transformers work?
Park, N. and Kim, S · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with CLIP latents
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Later among the works it cites.
When does sgd favor flat minima? a quantitative characterization via linear stability
Wu, L., Wang, M., and Su, W · 2022
Later among the works it cites.
Surrogate gap minimization improves sharpness-aware training
Zhuang, J., Gong, B., Yuan, L., Cui, Y., Adam, H., Dvornek, N. C., sekhar tatikonda, s Duncan, J., and Liu, T · 2022
Later among the works it cites.
SGD with large step sizes learns sparse features
Andriushchenko, M., Varre, A., Pillaud-Vivien, L., and Flammarion, N · 2023
Closest in time.
Self-stabilization: The implicit bias of gradient descent at the edge of stability
Damian, A., Nichani, E., and Lee, J. D · 2023
Closest in time.