Fetching the paper…
Reading the bibliography…
How to train deep neural networks (DNNs) to generalize well is a central concern in deep learning, especially for severely overparameterized networks nowadays.
A simple weight decay can improve generalization
Krogh, A. and Hertz, J. A · 1991
Earlier work this paper cites.
Fast exact multiplication by the hessian
Pearlmutter, B. A · 1994
Earlier work this paper cites.
Flat minima
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Penalty functions
Smith, A. E., Coit, D. W., Baeck, T., Fogel, D., and Michalewicz, Z · 1997
Earlier work this paper cites.
Regularization of neural networks using dropconnect
Wan, L., Zeiler, M. D., Zhang, S., LeCun, Y., and Fergus, R · 2013
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G. E., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Path-sgd: Path-normalized optimization in deep neural networks
Neyshabur, B., Salakhutdinov, R., and Srebro, N · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2015
Earlier work this paper cites.
Ba, L. J., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Deep Learning
Goodfellow, I. J., Bengio, Y., and Courville, A. C · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Wide residual networks
Zagoruyko, S. and Komodakis, N · 2016
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J. T., Sagun, L., and Zecchina, R · 2017
Cited alongside, same era.
Improved regularization of convolutional neural networks with cutout
Devries, T. and Taylor, G. W · 2017
Cited alongside, same era.
Sharp minima can generalize for deep nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y · 2017
Cited alongside, same era.
Shake-shake regularization of 3-branch residual networks
Gastaldi, X · 2017
Cited alongside, same era.
Lipschitz regularity of deep neural networks: analysis and efficient estimation
Virmaux, A. and Scaman, K · 2018
Later among the works it cites.
How SGD selects the global minima in over-parameterized learning: A dynamical stability perspective
Wu, L., Ma, C., and E, W · 2018
Later among the works it cites.
Group normalization
Wu, Y. and He, K · 2018
Later among the works it cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Later among the works it cites.
Shakedrop regularization for deep residual learning
Yamada, Y., Iwamura, M., Akiba, T., and Kise, K · 2019
Later among the works it cites.
Cutmix: Regularization strategy to train strong classifiers with localizable features
Yun, S., Han, D., Chun, S., Oh, S. J., Yoo, Y., and Choe, J · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Goyal, P., Dollár, P., Girshick, R. B., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Cited alongside, same era.
Deep pyramidal residual networks
Han, D., Kim, J., and Kim, J · 2017
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2017
Cited alongside, same era.
Exploring generalization in deep learning
Neyshabur, B., Bhojanapalli, S., McAllester, D., and Srebro, N · 2017
Cited alongside, same era.
Spectral norm regularization for improving the generalizability of deep learning
Yoshida, Y. and Miyato, T · 2017
Cited alongside, same era.
Autoaugment: Learning augmentation policies from data
Cubuk, E. D., Zoph, B., Mané, D., Vasudevan, V., and Le, Q. V · 2018
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Later among the works it cites.
Sharpness-aware minimization for efficiently improving generalization
Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B · 2021
Later among the works it cites.
ASAM: adaptive sharpness-aware minimization for scale-invariant learning of deep neural networks
Kwon, J., Kim, J., Park, H., and Choi, I. K · 2021
Later among the works it cites.
A diffusion theory for deep learning dynamics: Stochastic gradient descent exponentially favors flat minima
Xie, Z., Sato, I., and Sugiyama, M · 2021
Later among the works it cites.
Regularizing neural networks via adversarial model perturbation
Zheng, Y., Zhang, R., and Mao, Y · 2021
Later among the works it cites.