Exploring generalization in deep learning
B. Neyshabur, S. Bhojanapalli, D. McAllester, and N. Srebro · 2017
Later among the works it cites.
Empirical analysis of the Hessian of over-parametrized neural networks
Original
L. Sagun, U. Evci, V. U. Guney, Y. Dauphin, and L. Bottou · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Later among the works it cites.
Stronger generalization bounds for deep nets via a compression approach
S. Arora, R. Ge, B. Neyshabur, and Y. Zhang · 2018
Later among the works it cites.
How many samples are needed to estimate a convolutional neural network?
S. S. Du, Y. Wang, X. Zhai, S. Balakrishnan, R. R. Salakhutdinov, and A. Singh · 2018
Later among the works it cites.
Size-independent sample complexity of neural networks
N. Golowich, A. Rakhlin, and O. Shamir · 2018
Later among the works it cites.
Implicit bias of gradient descent on linear convolutional networks
S. Gunasekar, J. D. Lee, D. Soudry, and N. Srebro · 2018
Later among the works it cites.
Deep linear networks with arbitrary loss: All local minima are global
T. Laurent and J. Brecht · 2018
Later among the works it cites.
On tighter generalization bound for deep neural networks: Cnns, resnets, and beyond
Original
X. Li, J. Lu, Z. Wang, J. Haupt, and T. Zhao · 2018
Later among the works it cites.
Foundations of machine learning
M. Mohri, A. Rostamizadeh, and A. Talwalkar · 2018
Later among the works it cites.
A PAC-Bayesian approach to spectrally-normalized margin bounds for neural networks
B. Neyshabur, S. Bhojanapalli, and N. Srebro · 2018
Later among the works it cites.
The Full Spectrum of Deepnet Hessians at Scale: Dynamics with SGD Training and Sample Size
Original
V. Papyan · 2018
Later among the works it cites.
A Bayesian perspective on generalization and stochastic gradient descent
S. L. Smith and Q. V. Le · 2018
Later among the works it cites.
The implicit bias of gradient descent on separable data
D. Soudry, E. Hoffer, M. S. Nacson, S. Gunasekar, and N. Srebro · 2018
Later among the works it cites.
High-dimensional probability: An introduction with applications in data science , volume 47
R. Vershynin · 2018
Later among the works it cites.
Understanding generalization and optimization performance of deep CNNs
P. Zhou and J. Feng · 2018
Later among the works it cites.
Nearly-tight VC-dimension and Pseudodimension Bounds for Piecewise Linear Neural Networks
P. L. Bartlett, N. Harvey, C. Liaw, and A. Mehrabian · 2019
Later among the works it cites.
A generalization theory of gradient descent for learning over-parameterized deep ReLU networks
Original
Y. Cao and Q. Gu · 2019
Later among the works it cites.
Entropy-SGD: Biasing gradient descent into wide valleys
P. Chaudhari, A. Choromanska, S. Soatto, Y. LeCun, C. Baldassi, C. Borgs, J. Chayes, L. Sagun, and R. Zecchina · 2019
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
S. S. Du, X. Zhai, B. Poczos, and A. Singh · 2019
Later among the works it cites.
Algorithm-dependent generalization bounds for overparameterized deep residual networks
S. Frei, Y. Cao, and Q. Gu · 2019
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
J. Lee, L. Xiao, S. Schoenholz, Y. Bahri, R. Novak, J. Sohl-Dickstein, and J. Pennington · 2019
Later among the works it cites.
Size-free generalization bounds for convolutional neural networks
Original
P. M. Long and H. Sedghi · 2019
Later among the works it cites.
Measurements of three-level hierarchical structure in the outliers in the spectrum of deepnet Hessians
V. Papyan · 2019
Later among the works it cites.
Fast-rate PAC-Bayes generalization bounds via shifted Rademacher processes
J. Yang, S. Sun, and D. M. Roy · 2019
Later among the works it cites.
Hessian based analysis of SGD for deep nets: Dynamics and generalization
X. Li, Q. Gu, Y. Zhou, T. Chen, and A. Banerjee · 2020
Closest in time.
In defense of uniform convergence: Generalization via derandomization with an application to interpolating predictors
J. Negrea, G. K. Dziugaite, and D. M. Roy · 2020
Closest in time.