Fetching the paper…
Reading the bibliography…
Modern machine learning usually involves predictors in the overparameterised setting (number of trained parameters greater than dataset size), and their training yields not only good performance on training data, but also good generalisation capacity.
I I -Divergence Geometry of Probability Distributions and Minimization Problems
I. Csiszár · 1975
Earlier work this paper cites.
Logarithmic Sobolev Inequalities
Leonard Gross · 1975
Earlier work this paper cites.
On extensions of the Brunn-Minkowski and Prékopa-Leindler theorems, including inequalities for log concave functions, and with an application to the diffusion equation
Herm Brascamp and Elliott Lieb · 1976
Earlier work this paper cites.
Asymptotic evaluation of certain Markov process expectations for large time—III
M. D. Donsker and S. R. S. Varadhan · 1976
Earlier work this paper cites.
A Generalized Poincaré Inequality for Gaussian Measures
William Beckner · 1989
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
A PAC Analysis of a Bayesian Estimator
John Shawe-Taylor and Robert C. Williamson · 1997
Earlier work this paper cites.
The MNIST database of handwritten digits
Yann LeCun · 1998
Earlier work this paper cites.
Some PAC-Bayesian Theorems
David A. McAllester · 1999
Earlier work this paper cites.
Sur les inégalités de Sobolev logarithmiques
Cécile Ané, Sébastien Blachère, Djalil Chafaï, Pierre Fougères, Ivan Gentil, Florent Malrieu, Cyril Roberto, and Grégory Scheffer · 2000
Earlier work this paper cites.
Generalization of an Inequality by Talagrand and Links with the Logarithmic Sobolev Inequality
F. Otto and C. Villani · 2000
Earlier work this paper cites.
Lectures on Logarithmic Sobolev Inequalities
Alice Guionnet and Bogusław Zegarlinksi · 2003
Earlier work this paper cites.
Entropies, convexity, and functional inequalities, On Φ \Phi -entropies and Φ \Phi -Sobolev inequalities
Djalil Chafaï · 2004
Earlier work this paper cites.
Concentration of measure and logarithmic Sobolev inequalities
Michel Ledoux · 2006
Earlier work this paper cites.
Information-theoretic upper and lower bounds for statistical estimation
Tong Zhang · 2006
Earlier work this paper cites.
PAC-Bayesian supervised classification: the thermodynamics of statistical learning
Olivier Catoni · 2007
Earlier work this paper cites.
Optimal transport: old and new
Cédric Villani · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Stochastic First- and Zeroth-order Methods for Nonconvex Stochastic Programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
PAC-Bayes-Empirical-Bernstein Inequality
Ilya Tolstikhin and Yevgeny Seldin · 2013
Cited alongside, same era.
Mathematical statistics: basic ideas and selected topics
Peter Bickel and Kjell Doksum · 2015
Cited alongside, same era.
Striving for Simplicity: The All Convolutional Net
Jost Springenberg, Alexey Dosovitskiy, Thomas Brox, and Martin Riedmiller · 2015
Cited alongside, same era.
On the properties of variational approximations of Gibbs posteriors
Pierre Alquier, James Ridgway, and Nicolas Chopin · 2016
Cited alongside, same era.
Computing Nonvacuous Generalization Bounds for Deep (Stochastic) Neural Networks with Many More Parameters than Training Data
Gintare Dziugaite and Daniel Roy · 2017
Cited alongside, same era.
Sharpness-aware Minimization for Efficiently Improving Generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2021
Later among the works it cites.
Information-Theoretic Generalization Bounds for Stochastic Gradient Descent
Gergely Neu · 2021
Later among the works it cites.
Progress in Self-Certified Neural Networks
María Pérez-Ortiz, Omar Rivasplata, Emilio Parrado-Hernández, Benjamin Guedj, and John Shawe-Taylor · 2021
Later among the works it cites.
Relative Flatness and Generalization
Henning Petzka, Michael Kamp, Linara Adilova, Cristian Sminchisescu, and Mario Boley · 2021
Later among the works it cites.
Integral Probability Metrics PAC-Bayes Bounds
Ron Amit, Baruch Epstein, Shay Moran, and Ron Meir · 2022
Later among the works it cites.
On the Importance of Gradient Norm in PAC-Bayesian Bounds
Itai Gat, Yossi Adi, Alexander Schwing, and Tamir Hazan · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stanislaw Jastrzebski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos J. Storkey · 2017
Cited alongside, same era.
Exploring Generalization in Deep Learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nati Srebro · 2017
Cited alongside, same era.
Gradient Descent Only Converges to Minimizers: Non-Isolated Critical Points and Invariant Regions
Ioannis Panageas and Georgios Piliouras · 2017
Cited alongside, same era.
Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms, 2017
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Cited alongside, same era.
Towards Deep Learning Models Resistant to Adversarial Attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2018
Cited alongside, same era.
A Primer on PAC-Bayesian Learning
Benjamin Guedj · 2019
Cited alongside, same era.
Distribution-Dependent Analysis of Gibbs-ERM Principle
Ilja Kuzborskij, Nicolò Cesa-Bianchi, and Csaba Szepesvári · 2019
Cited alongside, same era.
Later among the works it cites.
Online PAC-Bayes Learning
Maxime Haddouche and Benjamin Guedj · 2022
Later among the works it cites.
High Probability Guarantees for Nonconvex Stochastic Gradient Descent with Heavy Tails
Shaojie Li and Yong Liu · 2022
Later among the works it cites.
Generalization Bounds via Convex Analysis
Gábor Lugosi and Gergely Neu · 2022
Later among the works it cites.
A Modern Look at the Relationship between Sharpness and Generalization
Maksym Andriushchenko, Francesco Croce, Maximilian Müller, Matthias Hein, and Nicolas Flammarion · 2023
Later among the works it cites.
A Unified Recipe for Deriving (Time-Uniform) PAC-Bayes Bounds
Ben Chugg, Hongjian Wang, and Aaditya Ramdas · 2023
Later among the works it cites.
Handbook of Convergence Theorems for (Stochastic) Gradient Methods
Guillaume Garrigos and Robert M. Gower · 2023
Later among the works it cites.
Generalization Bounds: Perspectives from Information Theory and PAC-Bayes
Fredrik Hellström, Giuseppe Durisi, Benjamin Guedj, and Maxim Raginsky · 2023
Later among the works it cites.
Online-to-PAC Conversions: Generalization Bounds via Regret Analysis
Gábor Lugosi and Gergely Neu · 2023
Later among the works it cites.
Sharpness Minimization Algorithms Do Not Only Minimize Sharpness To Achieve Better Generalization
Kaiyue Wen, Zhiyuan Li, and Tengyu Ma · 2023
Later among the works it cites.
Sharpness-Aware Minimization Revisited: Weighted Sharpness as a Regularization Term
Yun Yue, Jiadi Jiang, Zhiling Ye, Ning Gao, Yongchao Liu, and Ke Zhang · 2023
Later among the works it cites.
User-friendly Introduction to PAC-Bayes Bounds
Pierre Alquier · 2024
Closest in time.
Fantastic Generalization Measures are Nowhere to be Found
Michael Gastpar, Ido Nachum, Jonathan Shafer, and Thomas Weinberger · 2024
Closest in time.