Fetching the paper…
Reading the bibliography…
Since the celebrated works of Russo and Zou (2016,2019) and Xu and Raginsky (2017), it has been well known that the generalization error of supervised learning algorithms can be bounded in terms of the mutual information between their input and the output, given that the loss of any fixed hypothesis has a subgaussian tail.
Uniformly convex spaces
James A Clarkson · 1936
Earlier work this paper cites.
Limit distributions for sums of independent random variables
Boris V Gnedenko and Andrey N Kolmogorov · 1954
Earlier work this paper cites.
On measures of entropy and information
Alfréd Rényi · 1961
Earlier work this paper cites.
Eine informationstheoretische ungleichung und ihre anwendung auf beweis der ergodizitaet von markoffschen ketten
Imre Csiszár · 1964
Earlier work this paper cites.
A course on empirical processes
Richard M Dudley · 1984
Earlier work this paper cites.
Real analysis , volume 32
Halsey Lawrence Royden and Patrick Fitzpatrick · 1988
Earlier work this paper cites.
Sharp uniform convexity and smoothness inequalities for trace norms
Keith Ball, Eric A Carlen, and Elliott H Lieb · 1994
Earlier work this paper cites.
The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network
Peter L Bartlett · 1998
Earlier work this paper cites.
Entropy numbers of linear function classes
Robert C Williamson, Alexander J Smola, and Bernhard Schölkopf · 2000
Earlier work this paper cites.
Fundamentals of Convex Analysis
Jean-Baptiste Hiriart-Urruty and Claude Lemaréchal · 2001
Earlier work this paper cites.
Bounds for averaging classifiers
John Langford and Matthias Seeger · 2001
Earlier work this paper cites.
PAC-Bayesian generalisation error bounds for gaussian process classification
Matthias Seeger · 2002
Earlier work this paper cites.
Covering number bounds of certain regularized linear function classes
Tong Zhang · 2002
Earlier work this paper cites.
Convex analysis in general vector spaces
Constantin Zǎlinescu · 2002
Earlier work this paper cites.
Topics in optimal transportation , volume 58
Cédric Villani · 2003
Earlier work this paper cites.
On the complexity of linear prediction: Risk bounds, margin bounds, and regularization
Sham M Kakade, Karthik Sridharan, and Ambuj Tewari · 2008
Cited alongside, same era.
On Rényi and Tsallis entropies and divergences for exponential families
Frank Nielsen and Richard Nock · 2011
Cited alongside, same era.
Online learning and online convex optimization
Shai Shalev-Shwartz · 2012
Cited alongside, same era.
Rényi divergence and Kullback–Leibler divergence
Tim van Erven and Peter Harremoës · 2014
Cited alongside, same era.
Fast rates in statistical and online learning
Tim van Erven, Peter D Grünwald, Nishant A Mehta, Mark D Reid, and Robert C Williamson · 2015
Cited alongside, same era.
PAC-Bayesian bounds based on the Rényi divergence
Luc Bégin, Pascal Germain, François Laviolette, and Jean-Francis Roy · 2016
A modern introduction to online learning
Francesco Orabona · 2019
Later among the works it cites.
How much does your data exploration overfit? controlling bias via information usage
Daniel Russo and James Zou · 2019
Later among the works it cites.
An information-theoretic view of generalization via Wasserstein distance
Hao Wang, Mario Diaz, José Cândido S Santos Filho, and Flavio P Calmon · 2019
Later among the works it cites.
Connections between mirror descent, Thompson sampling and the information ratio
Julian Zimmert and Tor Lattimore · 2019
Later among the works it cites.
Tightening mutual information-based bounds on generalization error
Yuheng Bu, Shaofeng Zou, and Venugopal V Veeravalli · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Introduction to online convex optimization
Elad Hazan · 2016
Cited alongside, same era.
Controlling bias in adaptive data analysis using information theory
Daniel Russo and James Zou · 2016
Cited alongside, same era.
f f -divergence inequalities
Igal Sason and Sergio Verdú · 2016
Cited alongside, same era.
Information-theoretic analysis of generalization capability of learning algorithms
Aolin Xu and Maxim Raginsky · 2017
Cited alongside, same era.
Simpler PAC-Bayesian bounds for hostile data
Pierre Alquier and Benjamin Guedj · 2018
Cited alongside, same era.
Learners that use little information
Raef Bassily, Shay Moran, Ido Nachum, Jonathan Shafer, and Amir Yehudayoff · 2018
Cited alongside, same era.
Mahdi Haghifam, Jeffrey Negrea, Ashish Khisti, Daniel M Roy, and Gintare Karolina Dziugaite · 2020
Later among the works it cites.
Generalization error bounds via m m th central moments of the information density
Fredrik Hellström and Giuseppe Durisi · 2020
Later among the works it cites.
Strongly convex divergences
James Melbourne · 2020
Later among the works it cites.
Reasoning about generalization via conditional mutual information
Thomas Steinke and Lydia Zakynthinou · 2020
Later among the works it cites.
User-friendly introduction to PAC-Bayes bounds
Pierre Alquier · 2021
Later among the works it cites.
Generalization error bounds via Rényi-, f f -divergences and maximal leakage
Amedeo Roberto Esposito, Michael Gastpar, and Ibrahim Issa · 2021
Later among the works it cites.
Towards a unified information-theoretic framework for generalization
Mahdi Haghifam, Gintare Karolina Dziugaite, Shay Moran, and Dan Roy · 2021
Later among the works it cites.
Information-theoretic generalization bounds for stochastic gradient descent
Gergely Neu, Gintare Karolina Dziugaite, Mahdi Haghifam, and Daniel M Roy · 2021
Later among the works it cites.
Tighter expected generalization error bounds via Wasserstein distance
Borja Rodríguez-Gálvez, Germán Bassi, Ragnar Thobaben, and Mikael Skoglund · 2021
Later among the works it cites.