Fetching the paper…
Reading the bibliography…
In an attempt to better understand generalization in deep learning, we study several possible explanations.
On the uniform convergence of relative frequencies of events to their probabilities
Vladimir~N Vapnik and A~Ya Chervonenkis · 1971
Earlier work this paper cites.
The best constants in the khintchine inequality
Uffe Haagerup · 1981
Earlier work this paper cites.
Occam's razor
Anselm Blumer, Andrzej Ehrenfeucht, David Haussler, and Manfred~K Warmuth · 1987
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White · 1989
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitchell~P Marcus, Mary~Ann Marcinkiewicz, and Beatrice Santorini · 1993
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, and Paolo Frasconi · 1994
Earlier work this paper cites.
Efficient agnostic learning of neural networks with bounded fan-in
Wee~Sun Lee, Peter~L Bartlett, and Robert~C Williamson · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Almost linear vc dimension bounds for piecewise polynomial networks
Peter~L Bartlett, Vitaly Maiorov, and Ron Meir · 1998
Earlier work this paper cites.
Some PAC-Bayesian theorems
David~A McAllester · 1998
Earlier work this paper cites.
The vanishing gradient problem during learning recurrent neural nets and problem solutions
Sepp Hochreiter · 1998
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
Efficient backprop
Yann Le Cun, Léon Bottou, Genevieve~B. Orr, and Klaus-Robert Müller · 1998
Earlier work this paper cites.
PAC-Bayesian model averaging
David~A McAllester · 1999
Earlier work this paper cites.
(not) bounding the true error
John Langford and Rich Caruana · 2001
Earlier work this paper cites.
A rank minimization heuristic with application to minimum order system approximation
Maryam Fazel, Haitham Hindi, and Stephen~P. Boyd · 2001
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter~L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Empirical margin distributions and bounding the generalization error of combined classifiers
Vladimir Koltchinskii and Dmitry Panchenko · 2002
Earlier work this paper cites.
Simplified pac-bayesian margin bounds
David McAllester · 2003
Earlier work this paper cites.
Pac-bayes & margins
John Langford and John Shawe-Taylor · 2003
Earlier work this paper cites.
Weighted low-rank approximations
Nathan Srebro and Tommi~S. Jaakkola · 2003
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter~L. Bartlett and Shahar Mendelson · 2003
Earlier work this paper cites.
Distance-based classification with lipschitz functions
Ulrike~von Luxburg and Olivier Bousquet · 2004
Earlier work this paper cites.
Maximum-margin matrix factorization
Nathan Srebro, Jason Rennie, and Tommi~S. Jaakkola · 2004
Earlier work this paper cites.
Fast maximum margin matrix factorization for collaborative prediction
Jasson~DM Rennie and Nathan Srebro · 2005
Earlier work this paper cites.
Convex neural networks
Yoshua Bengio, Nicolas~L. Roux, Pascal Vincent, Olivier Delalleau, and Patrice Marcotte · 2005
Earlier work this paper cites.
Rank, trace-norm and max-norm
Nathan Srebro and Adi Shraibman · 2005
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
Geoffrey~E Hinton, Simon Osindero, and Yee-Whye Teh · 2006
Earlier work this paper cites.
Introduction to the Theory of Computation
Michael Sipser · 2006
Cited alongside, same era.
Cryptographic hardness for learning intersections of halfspaces
Adam R Klivansand Alexander~A Sherstov · 2006
Cited alongside, same era.
Computational enhancements in low-rank semidefinite programming
Samuel Burer and Changhui Choi · 2006
Cited alongside, same era.
Cryptographic hardness for learning intersections of halfspaces
Adam~R Klivans and Alexander~A Sherstov · 2006
Cited alongside, same era.
Greedy layer-wise training of deep networks
Yoshua Bengio, Pascal Lamblin, Dan Popovici, Hugo Larochelle, et~al · 2007
Cited alongside, same era.
Topmoumoute online natural gradient algorithm
Nicolas~L Roux, Pierre-Antoine Manzagol, and Yoshua Bengio · 2008
Cited alongside, same era.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2014
Later among the works it cites.
A new perspective on learning linear separators with large lqlp margins
Maria-Florina Balcan and Christopher Berlind · 2014
Later among the works it cites.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Later among the works it cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Later among the works it cites.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Later among the works it cites.
Improving performance of recurrent neural network with relu nonlinearity
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kernel methods for deep learning
Youngmin Cho and Lawrence~K. Saul · 2009
Cited alongside, same era.
On the complexity of linear prediction: Risk bounds, margin bounds, and regularization
Sham~M Kakade, Karthik Sridharan, and AmbujTewari · 2009
Cited alongside, same era.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Cited alongside, same era.
Exploring strategies for training deep neural networks
Hugo Larochelle, Yoshua Bengio, Jér^ome Louradour, and Pascal Lamblin · 2009
Cited alongside, same era.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey~E Hinton · 2010
Cited alongside, same era.
Collaborative filtering in a non-uniform world: Learning with the weighted trace norm
Nathan Srebro and Ruslan Salakhutdinov · 2010
Cited alongside, same era.
Sachin~S. Talathi and Aniket Vartak · 2014
Later among the works it cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew~M Saxe, James~L McClelland, and Surya Ganguli · 2014
Later among the works it cites.
Revisiting natural gradient for deep networks
Razvan Pascanu and Yoshua Bengio · 2014
Later among the works it cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Later among the works it cites.
Riemannian metrics for neural networks ii: recurrent networks and learning symbolic data sequences
Yann Ollivier · 2015
Later among the works it cites.
A simple way to initialize recurrent networks of rectified linear units
Quoc~V Le, Navdeep Jaitly, and Geoffrey~E Hinton · 2015
Later among the works it cites.
Unitary evolution recurrent neural networks
Martin Arjovsky, Amar Shah, and Yoshua Bengio · 2015
Later among the works it cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Later among the works it cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Later among the works it cites.
Scaling up natural gradient by sparsely factorizing the inverse Fisher matrix
Roger Grosse and Ruslan Salakhudinov · 2015
Later among the works it cites.
Optimizing neural networks with Kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Later among the works it cites.
Guillaume Desjardins, Karen Simonyan, Razvan Pascanu, and Koray Kavukcuoglu · 2015
Later among the works it cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, and Yann LeCun · 2016
Later among the works it cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish~Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak~Peter Tang · 2016
Later among the works it cites.
Generalization error of invariant classifiers
Jure Sokolic, Raja Giryes, Guillermo Sapiro, and Miguel~RD Rodrigues · 2016
Later among the works it cites.
Regularizing RNNs by stabilizing activations
David Krueger and Roland Memisevic · 2016
Later among the works it cites.
The impact of the nonlinearity on the VC-dimension of a deep network
P.~L. Bartlett · 2017
Closest in time.
Nearly-tight vc-dimension bounds for piecewise linear neural networks
Nick Harvey, Chris Liaw, and Abbas Mehrabian · 2017
Closest in time.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Closest in time.
Gintare~Karolina Dziugaite and Daniel~M Roy · 2017
Closest in time.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nathan Srebro · 2017
Closest in time.
Spectrally-normalized margin bounds for neural networks
Peter Bartlett, Dylan~J Foster, and Matus Telgarsky · 2017
Closest in time.