Fetching the paper…
Reading the bibliography…
In practice it is often found that large over-parameterized neural networks generalize better than their smaller counterparts, an observation that appears to conflict with classical notions of function complexity, which typically favor smaller models.
Training a 3-node neural network is np-complete
Avrim Blum and Ronald L. Rivest · 1988
Earlier work this paper cites.
Learning representations by back-propagating errors
David E Rumelhart, Geoffrey E Hinton, Ronald J Williams, et al · 1988
Earlier work this paper cites.
Bayesian interpolation
David JC MacKay · 1991
Earlier work this paper cites.
Ockham’s razor and bayesian analysis
William H Jefferys and James O Berger · 1992
Earlier work this paper cites.
A simple weight decay can improve generalization
Anders Krogh and John A Hertz · 1992
Earlier work this paper cites.
A practical bayesian framework for backpropagation networks
David JC MacKay · 1992
Earlier work this paper cites.
Priors for infinite networks (tech. rep. no. crg-tr-94-1)
Radford M. Neal · 1994
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Occam’s razor
Carl E. Rasmussen and Zoubin Ghahramani · 2000
Earlier work this paper cites.
Derivative observations in gaussian process models of dynamic systems
Ercan Solak, Roderick Murray-Smith, William E Leithead, Douglas J Leith, and Carl E Rasmussen · 2003
Earlier work this paper cites.
A note on the evidence and bayesian occam’s razor, 2005
Iain Murray and Zoubin Ghahramani · 2005
Earlier work this paper cites.
A study of the effect of noise injection on the training of artificial neural networks
Yulei Jiang, Richard M Zur, Lorenzo L Pesce, and Karen Drukker · 2009
Earlier work this paper cites.
Convolutional deep belief networks on cifar-10
Alex Krizhevsky · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E Hinton · 2010
Earlier work this paper cites.
Contractive auto-encoders: Explicit invariance during feature extraction
Salah Rifai, Pascal Vincent, Xavier Muller, Xavier Glorot, and Yoshua Bengio · 2011
Earlier work this paper cites.
Neural networks for machine learning-lecture 6a-overview of mini-batch gradient descent, 2012
Geoffrey Hinton, Nitish Srivastava, and Kevin Swersky · 2012
Earlier work this paper cites.
Robustness and generalization
Huan Xu and Shie Mannor · 2012
Earlier work this paper cites.
On the number of response regions of deep feed forward networks with piece-wise linear activations
R. Pascanu, G. Montufar, and Y. Bengio · 2013
Cited alongside, same era.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Cited alongside, same era.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann Dauphin, Razvan Pascanu, Çaglar Gülçehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Cited alongside, same era.
Qualitatively characterizing neural network optimization problems
Ian J. Goodfellow and Oriol Vinyals · 2014
Cited alongside, same era.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy · 2014
Universal adversarial perturbations
Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, Omar Fawzi, and Pascal Frossard · 2016
Later among the works it cites.
The limitations of deep learning in adversarial settings
Nicolas Papernot, Patrick McDaniel, Somesh Jha, Matt Fredrikson, Z Berkay Celik, and Ananthram Swami · 2016
Later among the works it cites.
Exponential expressivity in deep neural networks through transient chaos
B. Poole, S. Lahiri, M. Raghu, J. Sohl-Dickstein, and S. Ganguli · 2016
Later among the works it cites.
On the Expressive Power of Deep Neural Networks
M. Raghu, B. Poole, J. Kleinberg, S. Ganguli, and J. Sohl-Dickstein · 2016
Later among the works it cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Towards deep neural network architectures robust to adversarial examples
Shixiang Gu and Luca Rigazio · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
On the Number of Linear Regions of Deep Neural Networks
G. Montúfar, R. Pascanu, K. Cho, and Y. Bengio · 2014
Cited alongside, same era.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Cited alongside, same era.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Cited alongside, same era.
Representation benefits of deep feedforward networks
Matus Telgarsky · 2015
Cited alongside, same era.
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, et al · 2016
Cited alongside, same era.
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Later among the works it cites.
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Later among the works it cites.
Google vizier: A service for black-box optimization
Daniel Golovin, Benjamin Solnik, Subhodeep Moitra, Greg Kochanski, John Karro, and D Sculley · 2017
Later among the works it cites.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nathan Srebro · 2017
Later among the works it cites.
Robust large margin deep neural networks
Jure Sokolic, Raja Giryes, Guillermo Sapiro, and Miguel RD Rodrigues · 2017
Later among the works it cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms, 2017
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Later among the works it cites.
Intriguing properties of adversarial examples
Ekin Dogus Cubuk, Barret Zoph, Samuel Stern Schoenholz, and Quoc V. Le · 2018
Closest in time.
Justin Gilmer, Luke Metz, Fartash Faghri, Sam Schoenholz, Maithra Raghu, Martin Wattenberg, and Ian Goodfellow · 2018
Closest in time.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Sam Schoenholz, Jeffrey Pennington, and Jascha Sohl-dickstein · 2018
Closest in time.
A PAC-bayesian approach to spectrally-normalized margin bounds for neural networks
Behnam Neyshabur, Srinadh Bhojanapalli, and Nathan Srebro · 2018
Closest in time.
Ensemble robustness and generalization of stochastic deep learning algorithms
Tom Zahavy, Bingyi Kang, Alex Sivak, Jiashi Feng, Huan Xu, and Shie Mannor · 2018
Closest in time.