Deep learning generalizes because the parameter-function map is biased towards simple functions
Original
Guillermo Valle-Pérez, Chico Q Camargo, and Ard A Louis · 2018
Later among the works it cites.
Energy–entropy competition and the effectiveness of stochastic gradient descent in machine learning
Yao Zhang, Andrew M Saxe, Madhu S Advani, and Alpha A Lee · 2018
Later among the works it cites.
Non-vacuous generalization bounds at the imagenet scale: a pac-bayesian compression approach
Wenda Zhou, Victor Veitch, Morgane Austern, Ryan P Adams, and Peter Orbanz · 2018
Later among the works it cites.
Learning transferable architectures for scalable image recognition
Barret Zoph, Vijay Vasudevan, Jonathon Shlens, and Quoc V Le · 2018
Later among the works it cites.
Finite size corrections for neural network gaussian processes
Original
Joseph M Antognini · 2019
Later among the works it cites.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Original
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Later among the works it cites.
Beyond linearization: On quadratic and higher-order approximation of wide neural networks
Yu Bai and Jason D Lee · 2019
Later among the works it cites.
Complexity, statistical risk, and metric entropy of deep nets using total path variation
Original
Andrew R Barron and Jason M Klusowski · 2019
Later among the works it cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Later among the works it cites.
Ordalia: Deep learning hyperparameter search via generalization error bounds extrapolation
Benedetto J Buratti and Eli Upfal · 2019
Later among the works it cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Yuan Cao and Quanquan Gu · 2019
Later among the works it cites.
Asymptotics of wide networks from feynman diagrams
Ethan Dyer and Guy Gur-Ari · 2019
Later among the works it cites.
High probability generalization bounds for uniformly stable algorithms with nearly optimal rate
Original
Vitaly Feldman and Jan Vondrák · 2019
Later among the works it cites.
A primer on pac-bayesian learning
Original
Benjamin Guedj · 2019
Later among the works it cites.
Dynamics of deep neural networks and neural tangent hierarchy
Original
Jiaoyang Huang and Horng-Tzer Yau · 2019
Later among the works it cites.
Fantastic generalization measures and where to find them
Original
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 2019
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Later among the works it cites.
On generalization error bounds of noisy gradient methods for non-convex learning
Original
Jian Li, Xuanyuan Luo, and Mingda Qiao · 2019
Later among the works it cites.
Uniform convergence may be unable to explain generalization in deep learning
Vaishnavh Nagarajan and J Zico Kolter · 2019
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Original
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2019
Later among the works it cites.
On the bias-variance tradeoff: Textbooks need an update
Original
Brady Neal · 2019
Later among the works it cites.
In defense of uniform convergence: Generalization via derandomization with an application to interpolating predictors
Original
Jeffrey Negrea, Gintare Karolina Dziugaite, and Daniel M Roy · 2019
Later among the works it cites.
The role of over-parametrization in generalization of neural networks
Behnam Neyshabur, Zhiyuan Li, Srinadh Bhojanapalli, Yann LeCun, and Nathan Srebro · 2019
Later among the works it cites.
Neural tangents: Fast and easy infinite neural networks in python, 2019
Roman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee, Alexander A. Alemi, Jascha Sohl-Dickstein, and Samuel S. Schoenholz · 2019
Later among the works it cites.
A constructive prediction of the generalization error across scales
Original
Jonathan S Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit · 2019
Later among the works it cites.
A primer on pac-bayesian learning
Benjamin Guedj John Shawe-Taylor · 2019
Later among the works it cites.
Asymptotic learning curves of kernel methods: empirical data vs teacher-student paradigm
Original
Stefano Spigler, Mario Geiger, and Matthieu Wyart · 2019
Later among the works it cites.
How noise affects the hessian spectrum in overparameterized neural networks
Original
Mingwei Wei and David J Schwab · 2019
Later among the works it cites.
Non-gaussian processes and neural networks at finite widths
Original
Sho Yaida · 2019
Later among the works it cites.
A fine-grained spectral perspective on neural networks
Original
Greg Yang and Hadi Salman · 2019
Later among the works it cites.
Understanding generalization error of sgd in nonconvex optimization
Yi Zhou, Yingbin Liang, and Huishuai Zhang · 2019
Later among the works it cites.
De-randomized pac-bayes margin bounds: Applications to non-convex and non-smooth predictors
Original
Arindam Banerjee, Tiancong Chen, and Yingxue Zhou · 2020
Closest in time.
Spectrum dependent learning curves in kernel regression and wide neural networks
Original
Blake Bordelon, Abdulkadir Canatar, and Cengiz Pehlevan · 2020
Closest in time.
A theory of universal learning, 2020
Olivier Bousquet, Steve Hanneke, Shay Moran, Ramon van Handel, and Amir Yehudayoff · 2020
Closest in time.
In search of robust measures of generalization
Original
Gintare Karolina Dziugaite, Alexandre Drouin, Brady Neal, Nitarshan Rajkumar, Ethan Caballero, Linbo Wang, Ioannis Mitliagkas, and Daniel M Roy · 2020
Closest in time.
On the marginal likelihood and cross-validation
Edwin Fong and CC Holmes · 2020
Closest in time.
Scaling laws for autoregressive generative modeling
Original
Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, Tom B Brown, Prafulla Dhariwal, Scott Gray, et al · 2020
Closest in time.
Scaling laws for neural language models
Original
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Closest in time.
Finite versus infinite neural networks: an empirical study
Jaehoon Lee, Samuel Schoenholz, Jeffrey Pennington, Ben Adlam, Lechao Xiao, Roman Novak, and Jascha Sohl-Dickstein · 2020
Closest in time.
Rethinking parameter counting in deep models: Effective dimensionality revisited
Original
Wesley J Maddox, Gregory Benton, and Andrew Gordon Wilson · 2020
Closest in time.
Is sgd a bayesian sampler? well, almost
Original
Chris Mingard, Guillermo Valle-Pérez, Joar Skalse, and Ard A Louis · 2020
Closest in time.
Towards nngp-guided neural architecture search
Original
Daniel S Park, Jaehoon Lee, Daiyi Peng, Yuan Cao, and Jascha Sohl-Dickstein · 2020
Closest in time.
Pac-bayes analysis beyond the usual bounds
Original
Omar Rivasplata, Ilja Kuzborskij, Csaba Szepesvári, and John Shawe-Taylor · 2020
Closest in time.
Bayesian deep learning and a probabilistic perspective of generalization
Original
Andrew Gordon Wilson and Pavel Izmailov · 2020
Closest in time.