Fetching the paper…
Reading the bibliography…
In Deep Learning, Stochastic Gradient Descent (SGD) is usually selected as a training method because of its efficiency; however, recently, a problem in SGD gains research interest: sharp minima in Deep Neural Networks (DNNs) have poor generalization; especially, large-batch SGD tends to converge to sharp minima.
C. S. Wallace and D. M. Boulton, “An information measure for classification,” The Computer Journal , vol. 11, no. 2, pp. 185–194, 1968
1968
Earlier work this paper cites.
J. Rissanen, “Modeling by shortest data description,” Automatica , vol. 14, no. 5, pp. 465–471, 1978
1978
Earlier work this paper cites.
J.-P. Bouchaud and A. Georges, “Anomalous diffusion in disordered media: statistical mechanisms, models and physical applications,” Physics reports , vol. 195, no. 4-5, pp. 127–293, 1990
1990
Earlier work this paper cites.
D. J. MacKay, “A practical bayesian framework for backpropagation networks,” Neural computation , vol. 4, no. 3, pp. 448–472, 1992
1992
Earlier work this paper cites.
G. Hinton and D. van Camp, “Keeping neural networks simple by minimising the description length of weights. 1993,” in Proceedings of COLT-93 , pp. 5–13
1993
Earlier work this paper cites.
R. H. Byrd, P. Lu, J. Nocedal, and C. Zhu, “A limited memory algorithm for bound constrained optimization,” SIAM Journal on Scientific Computing , vol. 16, no. 5, pp. 1190–1208, 1995
1995
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Flat minima,” Neural Computation , vol. 9, no. 1, pp. 1–42, 1997
1997
Earlier work this paper cites.
L. Bottou, “Online learning and stochastic approximations,” On-line learning in neural networks , vol. 17, no. 9, p. 142, 1998
1998
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998
1998
Earlier work this paper cites.
K. Ho, C.-s. Leung, and J. Sum, “On weight-noise-injection training,” in International Conference on Neural Information Processing . Springer, 2008, pp. 919–926
2008
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on . IEEE, 2009, pp. 248–255
2009
Earlier work this paper cites.
A. Krizhevsky and G. Hinton, “Learning multiple layers of features from tiny images,” 2009
2009
Earlier work this paper cites.
M. Welling and Y. W. Teh, “Bayesian learning via stochastic gradient langevin dynamics,” in Proceedings of the 28th International Conference on Machine Learning (ICML-11) , 2011, pp. 681–688
2011
Earlier work this paper cites.
R. M. Neal, Bayesian learning for neural networks . Springer Science & Business Media, 2012, vol. 118
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems , 2012, pp. 1097–1105
2012
Cited alongside, same era.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” arXiv:1312.6114 , 2013
2013
Cited alongside, same era.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,” The Journal of Machine Learning Research , vol. 15, no. 1, pp. 1929–1958, 2014
2014
Cited alongside, same era.
2014
Cited alongside, same era.
2017
Later among the works it cites.
2017
Later among the works it cites.
P. Chaudhari, A. Choromanska, S. Soatto, and Y. LeCun, “Entropy-sgd: Biasing gradient descent into wide valleys,” in International Conference on Learning Representations , 2017
2017
Later among the works it cites.
S. L. Smith and Q. V. Le, “A bayesian perspective on generalization and stochastic gradient descent,” in Proceedings of Second workshop on Bayesian Deep Learning (NIPS 2017) , 2017
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
2016
Cited alongside, same era.
H. Mobahi, “Training recurrent neural networks by diffusion,” arXiv:1601.04114 , 2016
2016
Cited alongside, same era.
2016
Cited alongside, same era.
Y. Gal and Z. Ghahramani, “Dropout as a bayesian approximation: Representing model uncertainty in deep learning,” in international conference on machine learning , 2016, pp. 1050–1059
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 770–778
2016
Cited alongside, same era.
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang, “On large-batch training for deep learning: Generalization gap and sharp minima,” in International Conference on Learning Representations , 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
L. Dinh, R. Pascanu, S. Bengio, and Y. Bengio, “Sharp minima can generalize for deep nets,” in International Conference on Machine Learning , 2017, pp. 1019–1028
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
S. L. Smith, P.-J. Kindermans, and Q. V. Le, “Don’t decay the learning rate, increase the batch size,” in International Conference on Learning Representations , 2018
2018
Closest in time.
S. Jastrzębski, Z. Kenton, D. Arpit, N. Ballas, A. Fischer, Y. Bengio, and A. Storkey, “Finding flatter minima with sgd,” 2018
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
G. Diamos, S. Sengupta, B. Catanzaro, M. Chrzanowski, A. Coates, E. Elsen, J. Engel, A. Hannun, and S. Satheesh, “Persistent rnns: Stashing recurrent weights on-chip,” in International Conference on Machine Learning , 2016, pp. 2024–2033
2033
Closest in time.