Fetching the paper…
Reading the bibliography…
This paper provides theoretical insights into why and how deep learning can generalize well, despite its large capacity, complexity, possible algorithmic instability, nonrobustness, and sharp minima, responding to an open question in the literature.
Universal approximation bounds for superpositions of a sigmoidal function
Andrew R Barron · 1993
Earlier work this paper cites.
Multilayer feedforward networks with a nonpolynomial activation function can approximate any function
Moshe Leshno, Vladimir Ya Lin, Allan Pinkus, and Shimon Schocken · 1993
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Structural risk minimization over data-dependent hierarchies
John Shawe-Taylor, Peter L Bartlett, Robert C Williamson, and Martin Anthony · 1998
Earlier work this paper cites.
Statistical learning theory , volume 1
Vladimir Vapnik · 1998
Earlier work this paper cites.
Rademacher processes and bounding the risk of function learning
Vladimir Koltchinskii and Dmitriy Panchenko · 2000
Earlier work this paper cites.
How to layer a directed acyclic graph
Patrick Healy and Nikola S Nikolov · 2001
Earlier work this paper cites.
Model selection and error estimation
Peter L Bartlett, Stéphane Boucheron, and Gábor Lugosi · 2002
Earlier work this paper cites.
Stability and generalization
Olivier Bousquet and André Elisseeff · 2002
Earlier work this paper cites.
Algorithmic luckiness
Ralf Herbrich and Robert C Williamson · 2002
Earlier work this paper cites.
Empirical margin distributions and bounding the generalization error of combined classifiers
Vladimir Koltchinskii and Dmitry Panchenko · 2002
Earlier work this paper cites.
Learning theory: stability is sufficient for generalization and necessary and sufficient for consistency of empirical risk minimization
Sayan Mukherjee, Partha Niyogi, Tomaso Poggio, and Ryan Rifkin · 2006
Earlier work this paper cites.
Foundations of machine learning
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar · 2012
Earlier work this paper cites.
User-friendly tail bounds for sums of random matrices
Joel A Tropp · 2012
Earlier work this paper cites.
Robustness and generalization
Huan Xu and Shie Mannor · 2012
Earlier work this paper cites.
Regularization of neural networks using dropconnect
Li Wan, Matthew Zeiler, Sixin Zhang, Yann L Cun, and Rob Fergus · 2013
Cited alongside, same era.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Cited alongside, same era.
On the number of linear regions of deep neural networks
Guido F Montufar, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
On the number of response regions of deep feed forward networks with piece-wise linear activations
Razvan Pascanu, Guido Montufar, and Yoshua Bengio · 2014
Cited alongside, same era.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Cited alongside, same era.
The loss surfaces of multilayer networks
Anna Choromanska, MIkael Henaff, Michael Mathieu, Gerard Ben Arous, and Yann LeCun · 2015
Benefits of depth in neural networks
Matus Telgarsky · 2016
Later among the works it cites.
Aggregated residual transformations for deep neural networks
Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He · 2016
Later among the works it cites.
A closer look at memorization in deep networks
Devansh Arpit, Stanislaw Jastrzebski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, et al · 2017
Closest in time.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Closest in time.
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Benjamin Recht, and Yoram Singer · 2015
Cited alongside, same era.
Bayesian optimization with exponential convergence
Kenji Kawaguchi, Leslie Pack Kaelbling, and Tomás Lozano-Pérez · 2015
Cited alongside, same era.
Apac: Augmented pattern classification with neural networks
Ikuro Sato, Hiroki Nishimura, and Kensuke Yokoi · 2015
Cited alongside, same era.
An introduction to matrix concentration inequalities
Joel A Tropp et al · 2015
Cited alongside, same era.
Pengtao Xie, Yuntian Deng, and Eric Xing · 2015
Cited alongside, same era.
Unsupervised learning for physical interaction through video prediction
Chelsea Finn, Ian Goodfellow, and Sergey Levine · 2016
Cited alongside, same era.
Alon Gonen and Shai Shalev-Shwartz · 2017
Closest in time.
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Closest in time.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Closest in time.
Deep nets don’t learn via memorization
David Krueger, Nicolas Ballas, Stanislaw Jastrzebski, Devansh Arpit, Maxinder S Kanwal, Tegan Maharaj, Emmanuel Bengio, Asja Fischer, and Aaron Courville · 2017
Closest in time.
Data-dependent stability of stochastic gradient descent
Ilja Kuzborskij and Christoph Lampert · 2017
Closest in time.
Why and when can deep-but not shallow-networks avoid the curse of dimensionality: A review
Tomaso Poggio, Hrushikesh Mhaskar, Lorenzo Rosasco, Brando Miranda, and Qianli Liao · 2017
Closest in time.
Towards understanding generalization of deep learning: Perspective of loss landscapes
Lei Wu, Zhanxing Zhu, et al · 2017
Closest in time.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Closest in time.
Generalization in machine learning via analytical learning theory
Kenji Kawaguchi and Yoshua Bengio · 2018
Closest in time.