Fetching the paper…
Reading the bibliography…
Existing bounds on the generalization error of deep networks assume some form of smooth or bounded dependence on the input variable, falling short of investigating the mechanisms controlling such factors in practice.
Verification of forecasts expressed in terms of probability
Glenn W Brier · 1950
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
A stochastic estimator of the trace of the influence matrix for laplacian smoothing splines
Michael F Hutchinson · 1990
Earlier work this paper cites.
Neural Networks and the Bias/Variance Dilemma
Stuart Geman, Elie Bienenstock, and René Doursat · 1992
Earlier work this paper cites.
Fast exact multiplication by the hessian
Barak A Pearlmutter · 1994
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Wit3: Web inventory of transcribed and translated talks
Mauro Cettolo, Christian Girardi, and Marcello Federico · 2012
Earlier work this paper cites.
Results of the wmt14 metrics shared task
Matouš Macháček and Ondřej Bojar · 2014
Earlier work this paper cites.
On the number of linear regions of deep neural networks
Guido F Montufar, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Earlier work this paper cites.
An exploration of softmax alternatives belonging to the spherical loss family
Alexandre De Brebisson and Pascal Vincent · 2016
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Ben Recht, and Yoram Singer · 2016
Earlier work this paper cites.
Benefits of depth in neural networks
Matus Telgarsky · 2016
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Earlier work this paper cites.
On the expressive power of deep neural networks
Maithra Raghu, Ben Poole, Jon Kleinberg, Surya Ganguli, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
On the optimization of deep networks: Implicit acceleration by overparameterization
Sanjeev Arora, Nadav Cohen, and Elad Hazan · 2018
Cited alongside, same era.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P Vetrov, and Andrew G Wilson · 2018
Cited alongside, same era.
Deterministic pac-bayesian generalization bounds for deep networks via generalizing noise-resilience
Vaishnavh Nagarajan and Zico Kolter · 2018
Cited alongside, same era.
The role of over-parametrization in generalization of neural networks
Behnam Neyshabur, Zhiyuan Li, Srinadh Bhojanapalli, Yann LeCun, and Nathan Srebro · 2018
Cited alongside, same era.
Sensitivity and generalization in neural networks: an empirical study
Benign overfitting in linear regression
Peter L. Bartlett, Philip M. Long, Gábor Lugosi, and Alexander Tsigler · 2020
Later among the works it cites.
The intriguing role of module criticality in the generalization of deep networks
Niladri S Chatterji, Behnam Neyshabur, and Hanie Sedghi · 2020
Later among the works it cites.
Exactly computing the local lipschitz constant of relu networks
Matt Jordan and Alexandros G Dimakis · 2020
Later among the works it cites.
Adversarial training is a form of data-dependent operator norm regularization
Kevin Roth, Yannic Kilcher, and Thomas Hofmann · 2020
Later among the works it cites.
On the interplay between noise and curvature and its effect on optimization and generalization
Valentin Thomas, Fabian Pedregosa, Bart van Merriënboer, Pierre-Antoine Manzagol, Yoshua Bengio, and Nicolas Le Roux · 2020
Later among the works it cites.
A universal law of robustness via isoperimetry
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Roman Novak, Yasaman Bahri, Daniel A Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Cited alongside, same era.
Lipschitz regularity of deep neural networks: analysis and efficient estimation
Aladin Virmaux and Kevin Scaman · 2018
Cited alongside, same era.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
Lei Wu and Chao Ma · 2018
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2018
Cited alongside, same era.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Cited alongside, same era.
Jamming transition as a paradigm to understand the loss landscape of deep neural networks
Mario Geiger, Stefano Spigler, Stéphane d’Ascoli, Levent Sagun, Marco Baity-Jesi, Giulio Biroli, and Matthieu Wyart · 2019
Cited alongside, same era.
Deep relu networks have surprisingly few activation patterns
Boris Hanin and David Rolnick · 2019
Cited alongside, same era.
Sébastien Bubeck and Mark Sellke · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2021
Later among the works it cites.
Regularisation of neural networks by enforcing lipschitz continuity
Henry Gouk, Eibe Frank, Bernhard Pfahringer, and Michael J Cree · 2021
Later among the works it cites.
Noise and fluctuation of finite learning rate stochastic gradient descent
Kangqiao Liu, Liu Ziyin, and Masahito Ueda · 2021
Later among the works it cites.
On linear stability of sgd and input-smoothness of neural networks
Chao Ma and Lexing Ying · 2021
Later among the works it cites.
Why neural networks find simple solutions: the many regularizers of geometric complexity
Benoit Dherin, Michael Munn, Mihaela Rosca, and David GT Barrett · 2022
Later among the works it cites.
Are all linear regions created equal?
Matteo Gamba, Adrian Chmielewski-Anders, Josephine Sullivan, Hossein Azizpour, and Mårten Björkman · 2022
Later among the works it cites.
On second-moment stability of discrete-time linear systems with general stochastic dynamics
Yohei Hosoe and Tomomichi Hagiwara · 2022
Later among the works it cites.
Robustness implies generalization via data-dependent generalization bounds
Kenji Kawaguchi, Zhun Deng, Kyle Luh, and Jiaoyang Huang · 2022
Later among the works it cites.
Power-law escape rate of SGD
Takashi Mori, Liu Ziyin, Kangqiao Liu, and Masahito Ueda · 2022
Later among the works it cites.
Strength of minibatch noise in SGD
Liu Ziyin, Kangqiao Liu, Takashi Mori, and Masahito Ueda · 2022
Later among the works it cites.