Fetching the paper…
Reading the bibliography…
It was empirically confirmed by Keskar et al.\cite{SharpMinima} that flatter minima generalize better.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Some pac-bayesian theorems
D. A. McAllesterl · 1998
Earlier work this paper cites.
Pac-bayesian model averaging
D. A. McAllesterl · 1999
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton · 2009
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
M. Hardt, B. Recht, and Y. Singer · 2015
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
O. V. Ian J Goodfellow and A. M. Saxe · 2015
Earlier work this paper cites.
Path-sgd: Path-normalized optimization in deep neural networks
B. Neyshabur, R. Salakhutdinov, and N. Srebro · 2015
Earlier work this paper cites.
Wide & deep learning for recommender systems
H.-T. Cheng, L. Koc, J. Harmsen, T. Shaked, T. Chandra, H. Aradhye, G. Anderson, G. Corrado, W. Chai, M. Ispir, et al · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2016
Cited alongside, same era.
Sharp minima can generalize for deep nets
L. Dinh, R. Pascanu, S. Bengio, and Y. Bengio · 2017
Cited alongside, same era.
Convolutional sequence to sequence learning
J. Gehring, M. Auli, D. Grangier, D. Yarats, and Y. N. Dauphin · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Later among the works it cites.
Deep & cross network for ad click predictions
R. Wang, B. Fu, G. Fu, and M. Wang · 2017
Later among the works it cites.
Three factors influencing minima in sgd
S. Jastrzlbski, Z. Kenton, D. Arpit, N. Ballas, A. Fischer, Y. Bengio, and A. Storkey · 2018
Later among the works it cites.
Visualizing the loss landscape of neural nets
H. Li, Z. Xu, G. Taylor, C. Studer, and T. Goldstein · 2018
Later among the works it cites.
G-sgd: Optimizing relu neural networks in its positively scale-invariant space
Q. Meng, S. Zheng, H. Zhang, W. Chen, Z.-M. Ma, and T.-Y. Liu · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Densely connected convolutional networks
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger · 2017
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2017
Cited alongside, same era.
Exploring generalization in deep learning
B. Neyshabur, S. Bhojanapalli, D. McAllester, and N. Srebro · 2017
Cited alongside, same era.
Exploring generalization in deep learning
B. Neyshabur, S. Bhojanapalli, D. McAllester, and N. Srebro · 2017
Cited alongside, same era.
A bayesian perspective on generalization and stochastic gradient descent
S. L. Smith and Q. V. Le · 2018
Later among the works it cites.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
L. Wu, C. Ma, and W. E · 2018
Later among the works it cites.
Capacity control of relu neural networks by basis-path norm
S. Zheng, Q. Meng, H. Zhang, W. Chen, N. Yu, and T.-Y. Liu · 2018
Later among the works it cites.