Fetching the paper…
Reading the bibliography…
This dissertation studies a fundamental open challenge in deep learning theory: why do deep networks generalize well even while being overparameterized, unregularized and fitting the training data to zero error? In the first part of the thesis, we will empirically study how training deep networks via stochastic gradient descent implicitly controls the networks' capacity.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J. Tibshirani · 1903
Earlier work this paper cites.
Verification of probabilistic predictions: A brief review
Allan H Murphy and Edward S Epstein · 1967
Earlier work this paper cites.
On the Uniform Convergence of Relative Frequencies of Events to Their Probabilities
V. N. Vapnik and A. Ya. Chervonenkis · 1971
Earlier work this paper cites.
A finite sample distribution-free performance bound for local discrimination rules
W. H. Rogers and T. J. Wagner · 1978
Earlier work this paper cites.
The well-calibrated bayesian
A Philip Dawid · 1982
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell B. Stinchcombe, and Halbert White · 1989
Earlier work this paper cites.
Keeping the neural networks simple by minimizing the description length of the weights
Geoffrey E. Hinton and Drew van Camp · 1993
Earlier work this paper cites.
Reflections after refereeing papers for nips
Leo Breiman · 1995
Earlier work this paper cites.
Support-vector networks
Corinna Cortes and Vladimir Vapnik · 1995
Earlier work this paper cites.
Bagging predictors
Leo Breiman · 1996
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Boosting the margin: A new explanation for the effectiveness of voting methods
Robert E. Schapire, Yoav Freund, Peter Barlett, and Wee Sun Lee · 1997
Earlier work this paper cites.
The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network
Peter L Bartlett · 1998
Earlier work this paper cites.
Pac-bayesian model averaging
David A. McAllester · 1999
Earlier work this paper cites.
(not) bounding the true error
John Langford and Rich Caruana · 2001
Earlier work this paper cites.
Obtaining calibrated probability estimates from decision trees and naive bayesian classifiers
Bianca Zadrozny and Charles Elkan · 2001
Earlier work this paper cites.
Stability and generalization
Olivier Bousquet and André Elisseeff · 2002
Earlier work this paper cites.
Pac-bayes & margins
John Langford and John Shawe-Taylor · 2002
Earlier work this paper cites.
Simplified pac-bayesian margin bounds
David McAllester · 2003
Earlier work this paper cites.
Classification vs regression in overparameterized regimes: Does the loss function matter?
Vidya Muthukumar, Adhyyan Narang, Vignesh Subramanian, Mikhail Belkin, Daniel J. Hsu, and Anant Sahai · 2005
Earlier work this paper cites.
Moment multicalibration for uncertainty estimation
Christopher Jung, Changhwa Lee, Mallesh M. Pai, Aaron Roth, and Rakesh Vohra · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Distributional generalization: A new kind of generalization
Preetum Nakkiran and Yamini Bansal · 2009
Earlier work this paper cites.
Learnability, stability and uniform convergence
Shai Shalev-Shwartz, Ohad Shamir, Nathan Srebro, and Karthik Sridharan · 2010
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Zeyuan Allen-Zhu and Yuanzhi Li · 2012
Earlier work this paper cites.
Neurips 2020 competition: Predicting generalization in deep learning
Yiding Jiang, Pierre Foret, Scott Yak, Daniel M Roy, Hossein Mobahi, Gintare Karolina Dziugaite, Samy Bengio, Suriya Gunasekar, Isabelle Guyon, and Behnam Neyshabur · 2012
Earlier work this paper cites.
Foundations of Machine Learning
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar · 2012
Earlier work this paper cites.
Representation based complexity measures for predicting generalization in deep learning
Parth Natekar and Manik Sharma · 2012
Earlier work this paper cites.
User-friendly tail bounds for sums of random matrices
Joel A. Tropp · 2012
Earlier work this paper cites.
Min Lin, Qiang Chen, and Shuicheng Yan · 2013
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
Obtaining well calibrated probabilities using bayesian binning
Mahdi Pakdaman Naeini, Gregory F. Cooper, and Milos Hauskrecht · 2015
Earlier work this paper cites.
Nachdiplom lecture: Statistics meets optimization, lecture 2
Martin Wainwright · 2015
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Ben Recht, and Yoram Singer · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Stability and generalization in structured prediction
Ben London, Bert Huang, and Lise Getoor · 2016
Cited alongside, same era.
A closer look at memorization in deep networks
Devansh Arpit, Stanislaw K. Jastrzebski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S. Kanwal, Tegan Maharaj, Asja Fischer, Aaron C. Courville, Yoshua Bengio, and Simon Lacoste-Julien · 2017
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Peter L. Bartlett, Dylan J. Foster, and Matus J. Telgarsky · 2017
Cited alongside, same era.
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data
Gintare Karolina Dziugaite and Daniel M. Roy · 2017
Cited alongside, same era.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger · 2017
Cited alongside, same era.
Nearly-tight vc-dimension bounds for piecewise linear neural networks
Uniform convergence may be unable to explain generalization
Vaishnavh Nagarajan and J. Zico Kolter · 2019
Later among the works it cites.
The role of over-parametrization in generalization of neural networks
Behnam Neyshabur, Zhiyuan Li, Srinadh Bhojanapalli, Yann LeCun, and Nathan Srebro · 2019
Later among the works it cites.
Measuring calibration in deep learning
Jeremy Nixon, Michael W Dusenberry, Linchuan Zhang, Ghassen Jerfel, and Dustin Tran · 2019
Later among the works it cites.
On the spectral bias of neural networks
Nasim Rahaman, Aristide Baratin, Devansh Arpit, Felix Draxler, Min Lin, Fred A. Hamprecht, Yoshua Bengio, and Aaron C. Courville · 2019
Later among the works it cites.
Identifying and understanding deep learning phenomena
Hanie Sedghi, Samy Bengio, Kenji Hata, Aleksander Madry, Ari Morcos, Behnam Neyshabur, Maithra Raghu, Ali Rahimi, Ludwig Schmidt, and Ying Xiao · 2019
Later among the works it cites.
Evaluating model calibration in classification
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nick Harvey, Christopher Liaw, and Abbas Mehrabian · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Cited alongside, same era.
Generalization in deep learning
Kenji Kawaguchi, Leslie Pack Kaelbling, and Yoshua Bengio · 2017
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Cited alongside, same era.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Cited alongside, same era.
Deeper, broader and artier domain generalization
Da Li, Yongxin Yang, Yi-Zhe Song, and Timothy M. Hospedales · 2017
Cited alongside, same era.
Generalization in deep networks: The role of distance from initialization
Vaishnavh Nagarajan and J. Zico Kolter · 2017
Cited alongside, same era.
Juozas Vaicenavicius, David Widmann, Carl R. Andersson, Fredrik Lindsten, Jacob Roll, and Thomas B. Schön · 2019
Later among the works it cites.
High-Dimensional Statistics: A Non-Asymptotic Viewpoint
Martin J. Wainwright · 2019
Later among the works it cites.
Calibration tests in multi-class classification: A unifying framework
David Widmann, Fredrik Lindsten, and Dave Zachariah · 2019
Later among the works it cites.
Non-vacuous generalization bounds at the imagenet scale: a PAC-bayesian compression approach
Wenda Zhou, Victor Veitch, Morgane Austern, Ryan P. Adams, and Peter Orbanz · 2019
Later among the works it cites.
Benign overfitting in linear regression
Peter L. Bartlett, Philip M. Long, Gábor Lugosi, and Alexander Tsigler · 2020
Later among the works it cites.
Two models of double descent for weak features
Mikhail Belkin, Daniel Hsu, and Ji Xu · 2020
Later among the works it cites.
A model of double descent for high-dimensional logistic regression
Zeyu Deng, Abla Kammoun, and Christos Thrampoulidis · 2020
Later among the works it cites.
In search of robust measures of generalization
Gintare Karolina Dziugaite, Alexandre Drouin, Brady Neal, Nitarshan Rajkumar, Ethan Caballero, Linbo Wang, Ioannis Mitliagkas, and Daniel M Roy · 2020
Later among the works it cites.
Distribution-free binary classification: prediction sets, confidence intervals and calibration
Chirag Gupta, Aleksandr Podkopaev, and Aaditya Ramdas · 2020
Later among the works it cites.
https://simons.berkeley.edu/news/research-vignette-generalization-and-interpolation , 2020
Daniel Hsu · 2020
Later among the works it cites.
On the multiple descent of minimum-norm interpolants and restricted lower isometry of kernels
Tengyuan Liang, Alexander Rakhlin, and Xiyu Zhai · 2020
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and double descent curve, 2020
Song Mei and Andrea Montanari · 2020
Later among the works it cites.
The generalization error of max-margin linear classifiers: High-dimensional asymptotics in the overparametrized regime, 2020
Andrea Montanari, Feng Ruan, Youngtak Sohn, and Jun Yan · 2020
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2020
Later among the works it cites.
In defense of uniform convergence: Generalization via derandomization with an application to interpolating predictors
Jeffrey Negrea, Gintare Karolina Dziugaite, and Daniel Roy · 2020
Later among the works it cites.
Why are bootstrapped deep ensembles not better?
Jeremy Nixon, Balaji Lakshminarayanan, and Dustin Tran · 2020
Later among the works it cites.
Sample complexity of uniform convergence for multicalibration
Eliran Shabat, Lee Cohen, and Yishay Mansour · 2020
Later among the works it cites.
Benign overfitting in ridge regression, 2020
A. Tsigler and P. L. Bartlett · 2020
Later among the works it cites.
On uniform convergence and low-norm interpolation learning
Lijia Zhou, Danica J. Sutherland, and Nati Srebro · 2020
Later among the works it cites.
Yu Bai, Song Mei, Huan Wang, and Caiming Xiong · 2021
Closest in time.
Risk bounds for over-parameterized maximum margin classification on sub-gaussian mixtures
Yuan Cao, Quanquan Gu, and Mikhail Belkin · 2021
Closest in time.
Finite-sample analysis of interpolating linear classifiers in the overparameterized regime
Niladri S. Chatterji and Philip M. Long · 2021
Closest in time.
RATT: leveraging unlabeled data to guarantee generalization
Saurabh Garg, Sivaraman Balakrishnan, J. Zico Kolter, and Zachary C. Lipton · 2021
Closest in time.
Linearized two-layers neural networks in high dimension
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2021
Closest in time.
Early-stopped neural networks are consistent
Ziwei Ji, Justin D. Li, and Matus Telgarsky · 2021
Closest in time.
Assessing generalization of sgd via disagreement, 2021
Yiding Jiang, Vaishnavh Nagarajan, Christina Baek, and J. Zico Kolter · 2021
Closest in time.
Towards an understanding of benign overfitting in neural networks
Zhu Li, Zhi-Hua Zhou, and Arthur Gretton · 2021
Closest in time.
Jishnu Mukhoti, Andreas Kirsch, Joost van Amersfoort, Philip H. S. Torr, and Yarin Gal · 2021
Closest in time.
Benign overfitting in binary classification of gaussian mixtures
Ke Wang and Christos Thrampoulidis · 2021
Closest in time.
Benign overfitting in multiclass classification: All roads lead to interpolation
Ke Wang, Vidya Muthukumar, and Christos Thrampoulidis · 2021
Closest in time.
Should ensemble members be calibrated?
Xixin Wu and Mark Gales · 2021
Closest in time.