Fetching the paper…
Reading the bibliography…
A recent line of works, initiated by Russo and Xu, has shown that the generalization error of a learning algorithm can be upper bounded by information measures.
R. Moddemeijer, “On estimation of entropy and mutual information of continuous distributions,” Signal processing , vol. 16, no. 3, pp. 233–248, 1989
1989
Earlier work this paper cites.
W. H. Wong and X. Shen, “Probability inequalities for likelihood ratios and convergence rates of sieve mles,” The Annals of Statistics , pp. 339–362, 1995
1995
Earlier work this paper cites.
V. Vapnik, The nature of statistical learning theory . Springer science & business media, 1999
1999
Earlier work this paper cites.
D. A. McAllester, “Some pac-bayesian theorems,” Machine Learning , vol. 37, no. 3, pp. 355–363, 1999
1999
Earlier work this paper cites.
O. Bousquet and A. Elisseeff, “Stability and Generalization,” Journal of Machine Learning Research , vol. 2, no. Mar, pp. 499–526, 2002
2002
Earlier work this paper cites.
2004
Earlier work this paper cites.
A. Kraskov, H. Stögbauer, and P. Grassberger, “Estimating mutual information,” Physical review E , vol. 69, no. 6, p. 066138, 2004
2004
Earlier work this paper cites.
P. L. Bartlett and S. Mendelson, “Empirical minimization,” Probability theory and related fields , vol. 135, no. 3, pp. 311–334, 2006
2006
Earlier work this paper cites.
P. L. Bartlett, M. I. Jordan, and J. D. McAuliffe, “Convexity, classification, and risk bounds,” Journal of the American Statistical Association , vol. 101, no. 473, pp. 138–156, 2006
2006
Earlier work this paper cites.
N. Cesa-Bianchi and G. Lugosi, Prediction, learning, and games . Cambridge university press, 2006
2006
Earlier work this paper cites.
H. Xu and S. Mannor, “Robustness and generalization,” Machine learning , vol. 86, no. 3, pp. 391–423, 2012
2012
Earlier work this paper cites.
S. Boucheron, G. Lugosi, and P. Massart, Concentration Inequalities: A Nonasymptotic Theory of Independence . OUP Oxford, Feb. 2013
2013
Cited alongside, same era.
T. Van Erven, P. Grunwald, N. A. Mehta, M. Reid, R. Williamson et al. , “Fast rates in statistical and online learning,” 2015
2015
Cited alongside, same era.
T. Koren and K. Levy, “Fast rates for exp-concave empirical risk minimization,” Advances in Neural Information Processing Systems , vol. 28, 2015
2015
Cited alongside, same era.
D. Russo and J. Zou, “Controlling bias in adaptive data analysis using information theory,” in Artificial Intelligence and Statistics . PMLR, 2016, pp. 1232–1240
2016
Cited alongside, same era.
M. Raginsky, A. Rakhlin, M. Tsao, Y. Wu, and A. Xu, “Information-theoretic analysis of stability and bias of learning algorithms,” in 2016 IEEE Information Theory Workshop (ITW) . IEEE, 2016, pp. 26–30
A. Asadi, E. Abbe, and S. Verdu, “Chaining Mutual Information and Tightening Generalization Bounds,” in Advances in Neural Information Processing Systems 31 , S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, Eds. Curran Associates, Inc., 2018, pp. 7234–7243
2018
Later among the works it cites.
2019
Later among the works it cites.
T. Steinke and L. Zakynthinou, “Reasoning about generalization via conditional mutual information,” in Conference on Learning Theory . PMLR, 2020, pp. 3437–3452
2020
Later among the works it cites.
Y. Bu, S. Zou, and V. V. Veeravalli, “Tightening mutual information-based bounds on generalization error,” IEEE Journal on Selected Areas in Information Theory , vol. 1, no. 1, pp. 121–130, 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
S. Hanneke, “Refined error bounds for several learning algorithms,” The Journal of Machine Learning Research , vol. 17, no. 1, pp. 4667–4721, 2016
2016
Cited alongside, same era.
A. Xu and M. Raginsky, “Information-theoretic analysis of generalization capability of learning algorithms,” in Proceedings of the 31st International Conference on Neural Information Processing Systems , 2017, pp. 2521–2530
2017
Cited alongside, same era.
N. Mehta, “Fast rates with high probability in exp-concave statistical learning,” in Artificial Intelligence and Statistics . PMLR, 2017, pp. 1085–1093
2017
Cited alongside, same era.
J. Jiao, Y. Han, and T. Weissman, “Dependence measures bounding the exploration bias for general measurements,” in 2017 IEEE International Symposium on Information Theory (ISIT) . IEEE, 2017, pp. 1475–1479
2017
Cited alongside, same era.
W. Gao, S. Kannan, S. Oh, and P. Viswanath, “Estimating mutual information for discrete-continuous mixtures,” Advances in neural information processing systems , vol. 30, 2017
2017
Cited alongside, same era.
P. D. Grünwald and N. A. Mehta, “Fast rates for general unbounded loss functions: From erm to generalized bayes.” J. Mach. Learn. Res. , vol. 21, pp. 56–1, 2020
2020
Later among the works it cites.
J. Zhu, “Semi-supervised learning: the case when unlabeled data is equally useful,” in Conference on Uncertainty in Artificial Intelligence . PMLR, 2020, pp. 709–718
2020
Later among the works it cites.
H. Hafez-Kolahi, Z. Golgooni, S. Kasaei, and M. Soleymani, “Conditioning and processing: Techniques to improve information-theoretic generalization bounds,” Advances in Neural Information Processing Systems , vol. 33, pp. 16 457–16 467, 2020
2020
Later among the works it cites.
2021
Later among the works it cites.
G. Aminian, Y. Bu, L. Toni, M. Rodrigues, and G. Wornell, “An exact characterization of the generalization error for the gibbs algorithm,” Advances in Neural Information Processing Systems , vol. 34, 2021
2021
Later among the works it cites.
R. Zhou, C. Tian, and T. Liu, “Individually conditional individual mutual information bound on generalization error,” IEEE Transactions on Information Theory , 2022
2022
Closest in time.