Fetching the paper…
Reading the bibliography…
While it has not yet been proven, empirical evidence suggests that model generalization is related to local properties of the optima which can be described via the Hessian.
Sharp Minima Can Generalize For Deep Nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 1938
Earlier work this paper cites.
Some pac-bayesian theorems
David A. McAllester · 1998
Earlier work this paper cites.
Pac-bayesian model averaging
David A. McAllester · 1999
Earlier work this paper cites.
(Not) Bounding the True Error
John Langford and Rich Caruana · 2001
Earlier work this paper cites.
Pac-bayes & margins
John Langford and John Shawe-Taylor · 2002
Earlier work this paper cites.
Simplified pac-bayesian margin bounds
David Mcallester · 2003
Earlier work this paper cites.
Cubic regularization of newton method and its global performance
Yurii Nesterov and B. T. Polyak · 2006
Earlier work this paper cites.
Acoustic modeling using deep belief networks
A. Mohamed, G. E. Dahl, and G. Hinton · 2011
Earlier work this paper cites.
Pac-bayesian inequalities for martingales
Yevgeny Seldin, François Laviolette, Nicolò Cesa-Bianchi, John Shawe-Taylor, and Peter Auer · 2011
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition
Geoffrey Hinton, Li Deng, Dong Yu, George Dahl, Abdel rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara Sainath, and Brian Kingsbury · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Earlier work this paper cites.
Pac-bayesian analysis of supervised, unsupervised, and reinforcement learning, 2012
Y. Seldin, F. Laviolette, and J. Shawe-Taylor · 2012
Cited alongside, same era.
Generating sequences with recurrent neural networks
Alex Graves · 2013
Cited alongside, same era.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts · 2013
Cited alongside, same era.
Linear Coupling: An Ultimate Unification of Gradient and Mirror Descent
Zeyuan Allen-Zhu and Lorenzo Orecchia · 2014
Cited alongside, same era.
Spatial pyramid pooling in deep convolutional networks for visual recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2014
Gintare Karolina Dziugaite and Daniel M. Roy · 2017
Later among the works it cites.
Accurate, large minibatch SGD: training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross B. Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Later among the works it cites.
Nearly-tight VC-dimension bounds for piecewise linear neural networks
Nick Harvey, Christopher Liaw, and Abbas Mehrabian · 2017
Later among the works it cites.
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Later among the works it cites.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens van der Maaten, and Kilian Q Weinberger · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Understanding Machine Learning: From Theory to Algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer T. Chayes, Levent Sagun, and Riccardo Zecchina · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
Sergey Zagoruyko and Nikos Komodakis · 2016
Cited alongside, same era.
Cifar-10 (canadian institute for advanced research)
Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton
Cited in the paper.
Later among the works it cites.
Three Factors Influencing Minima in SGD
Stanisław Jastrzȩbski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Later among the works it cites.
A PAC-Bayesian Approach to Spectrally-Normalized Margin Bounds for Neural Networks
Behnam Neyshabur, Srinadh Bhojanapalli, and Nathan Srebro · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Later among the works it cites.
Fast and scalable bayesian deep learning by weight-perturbation in adam
Mohammad Emtiyaz Khan, Didrik Nielsen, Voot Tangkaratt, Wu Lin, Yarin Gal, and Akash Srivastava · 2018
Closest in time.
The natural language decathlon: Multitask learning as question answering
Bryan McCann, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher · 2018
Closest in time.
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma · 2018
Closest in time.