Fetching the paper…
Reading the bibliography…
One of the fundamental problems in machine learning is generalization.
Flat minima
Hochreiter, S. & Schmidhuber, J · 1997
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I. & Salakhutdinov, R · 2014
Earlier work this paper cites.
Deep learning
LeCun, Y., Bengio, Y. & Hinton, G · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S. & Sun, J · 2015
Earlier work this paper cites.
Norm-based capacity control in neural networks
Neyshabur, B., Tomioka, R. & Srebro, N · 2015
Earlier work this paper cites.
Deep learning , vol. 1 (MIT Press, 2016)
Goodfellow, I., Courville, A. & Bengio, Y · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S. & Sun, J · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Y. et al · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D. et al · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B. & Vinyals, O · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M. & Tang, P. T. P · 2017
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Dinh, L., Pascanu, R., Bengio, S. & Bengio, Y · 2017
Cited alongside, same era.
Empirical analysis of the hessian of over-parametrized neural networks (2017)
Sagun, L., Evci, U., Guney, V. U., Dauphin, Y. & Bottou, L · 2017
Cited alongside, same era.
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Lian, X. et al · 2017
Cited alongside, same era.
Densely connected convolutional networks
Huang, G., Liu, Z., Van Der Maaten, L. & Weinberger, K. Q · 2017
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
Chaudhari, P. et al · 2017
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Bartlett, P., Foster, D. J. & Telgarsky, M · 2017
Fantastic generalization measures and where to find them
Jiang, Y., Neyshabur, B., Mobahi, H., Krishnan, D. & Bengio, S · 2020
Later among the works it cites.
New insights and perspectives on the natural gradient method
Martens, J · 2020
Later among the works it cites.
Improved sample complexities for deep networks and robust classification via an all-layer margin
Wei, C. & Ma, T · 2020
Later among the works it cites.
Shaping the learning landscape in neural networks around wide flat minima
Baldassi, C., Pittorino, F. & Zecchina, R · 2020
Later among the works it cites.
Generalization in deep networks: The role of distance from initialization
Nagarajan, V. & Kolter, J. Z · 2020
Later among the works it cites.
Highly accurate protein structure prediction with alphafold
Jumper, J. et al · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Chaudhari, P. & Soatto, S · 2018
Cited alongside, same era.
Asynchronous decentralized parallel stochastic gradient descent
Lian, X., Zhang, W., Zhang, C. & Liu, J · 2018
Cited alongside, same era.
The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects
Zhu, Z., Wu, J., Yu, B., Wu, L. & Ma, J · 2019
Cited alongside, same era.
Time matters in regularizing deep networks: Weight decay and data augmentation affect early learning dynamics, matter little near convergence
Golatkar, A. S., Achille, A. & Soatto, S · 2019
Cited alongside, same era.
Towards understanding the role of over-parametrization in generalization of neural networks
Neyshabur, B., Li, Z., Bhojanapalli, S., LeCun, Y. & Srebro, N · 2019
Cited alongside, same era.
Later among the works it cites.
The inverse variance–flatness relation in stochastic gradient descent is critical for finding flat minima
Feng, Y. & Tu, Y · 2021
Later among the works it cites.
Loss landscape dependent self-adjusting learning rates in decentralized stochastic gradient descent
Zhang, W. et al · 2021
Later among the works it cites.
Sharpness-aware minimization for efficiently improving generalization
Foret, P., Kleiner, A., Mobahi, H. & Neyshabur, B · 2021
Later among the works it cites.
Phases of learning dynamics in artificial neural networks in the absence or presence of mislabeled data
Feng, Y. & Tu, Y · 2021
Later among the works it cites.
Does the data induce capacity control in deep learning?
Yang, R., Mao, J. & Chaudhari, P · 2022
Closest in time.