Fetching the paper…
Reading the bibliography…
Existing Rademacher complexity bounds for neural networks rely only on norm control of the weight matrices and depend exponentially on depth via a product of the matrix norms.
The sizes of compact subsets of hilbert space and continuity of gaussian processes
RM Dudley · 1967
Earlier work this paper cites.
Computational graphs and rounding error
Friedrich L Bauer · 1974
Earlier work this paper cites.
Vapnik-chervonenkis dimension of recurrent neural networks
Pascal Koiran and Eduardo D Sontag · 1997
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Dropout training as adaptive regularization
Stefan Wager, Sida Wang, and Percy S Liang · 2013
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2015
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Benjamin Recht, and Yoram Singer · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Regularizing rnns by stabilizing activations
David Krueger and Roland Memisevic · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Sergey Zagoruyko and Nikos Komodakis · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Cited alongside, same era.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
Alon Brutzkus, Amir Globerson, Eran Malach, and Shai Shalev-Shwartz · 2017
Cited alongside, same era.
Size-independent sample complexity of neural networks
Noah Golowich, Alexander Rakhlin, and Ohad Shamir · 2017
Regularizing by the variance of the activations’ sample-variances
Etai Littwin and Lior Wolf · 2018
Later among the works it cites.
Towards understanding the role of over-parametrization in generalization of neural networks
Behnam Neyshabur, Zhiyuan Li, Srinadh Bhojanapalli, Yann LeCun, and Nathan Srebro · 2018
Later among the works it cites.
Sensitivity and generalization in neural networks: an empirical study
Roman Novak, Yasaman Bahri, Daniel A Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Later among the works it cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Later among the works it cites.
On the margin theory of feedforward neural networks
Colin Wei, Jason D Lee, Qiang Liu, and Tengyu Ma · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Cited alongside, same era.
Self-normalizing neural networks
Günter Klambauer, Thomas Unterthiner, Andreas Mayr, and Sepp Hochreiter · 2017
Cited alongside, same era.
Robust large margin deep neural networks
Jure Sokolić, Raja Giryes, Guillermo Sapiro, and Miguel RD Rodrigues · 2017
Cited alongside, same era.
Stronger generalization bounds for deep nets via a compression approach
Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang · 2018
Cited alongside, same era.
Data-dependent pac-bayes priors via differential privacy
Gintare Karolina Dziugaite and Daniel M Roy · 2018
Cited alongside, same era.
Risk and parameter convergence of logistic regression
Ziwei Ji and Matus Telgarsky · 2018
Cited alongside, same era.
Later among the works it cites.
Yuxin Wu and Kaiming He · 2018
Later among the works it cites.
Stabilizing gradients for deep neural networks via efficient svd parameterization
Jiong Zhang, Qi Lei, and Inderjit S Dhillon · 2018
Later among the works it cites.
On generalization bounds of a family of recurrent neural networks, 2019
Minshuo Chen, Xingguo Li, and Tuo Zhao · 2019
Closest in time.
Deterministic PAC-bayesian generalization bounds for deep networks via generalizing noise-resilience
Vaishnavh Nagarajan and Zico Kolter · 2019
Closest in time.
Improved sample complexities for deep networks and robust classification via an all-layer margin
Colin Wei and Tengyu Ma · 2019
Closest in time.
Chain rule — Wikipedia, the free encyclopedia, 2019
Wikipedia contributors · 2019
Closest in time.
Residual learning without normalization via better initialization
Hongyi Zhang, Yann N. Dauphin, and Tengyu Ma · 2019
Closest in time.