Fetching the paper…
Reading the bibliography…
An information-theoretic upper bound on the generalization error of supervised learning algorithms is derived.
N. Littlestone and M. Warmuth, “Relating data compression and learnability,” Technical report, University of California, Santa Cruz
1986
Earlier work this paper cites.
D. A. McAllester, “Some PAC-Bayesian theorems,” Machine Learning
1999
Earlier work this paper cites.
O. Bousquet and A. Elisseeff, “Stability and generalization,” J. Mach. Learn. Res
2002
Earlier work this paper cites.
A. Kraskov, H. Stögbauer, and P. Grassberger, “Estimating mutual information,” Physical Review E
2004
Earlier work this paper cites.
S. Boucheron, O. Bousquet, and G. Lugosi, “Theory of classification: A survey of some recent advances,” ESAIM: Probability and Statistics
2005
Earlier work this paper cites.
A. Elisseeff, T. Evgeniou, and M. Pontil, “Stability of randomized learning algorithms,” J. Mach. Learn. Res
2005
Earlier work this paper cites.
Cambridge University Press, 2009
M. Anthony and P. L. Bartlett, Neural network learning: Theoretical foundations · 2009
Earlier work this paper cites.
S. Shalev-Shwartz, O. Shamir, N. Srebro, and K. Sridharan, “Learnability, stability and uniform convergence,” J. Mach. Learn. Res
2010
Earlier work this paper cites.
M. Welling and Y. Teh, “Bayesian learning via stochastic gradient Langevin dynamics,” in Proc. International Conference on Machine Learning (ICML)
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” in Proc. Advances in Neural Information Processing Systems (NIPS)
2012
Earlier work this paper cites.
Oxford University Press, 2013
S. Boucheron, G. Lugosi, and P. Massart, Concentration Inequalities: A Nonasymptotic Theory of Independence · 2013
Earlier work this paper cites.
Cambridge University Press, 2014
S. Shalev-Shwartz and S. Ben-David, Understanding machine learning: From theory to algorithms · 2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2015
Cited alongside, same era.
I. M. Alabdulmohsin, “Algorithmic stability and uniform generalization,” in Proc. Advances in Neural Information Processing Systems (NIPS)
2015
Cited alongside, same era.
MIT press Cambridge, 2016
I. Goodfellow, Y. Bengio, A. Courville, and Y. Bengio, Deep Learning · 2016
Cited alongside, same era.
D. Russo and J. Zou, “Controlling bias in adaptive data analysis using information theory,” in Proc. International Conference on Artifical Intelligence and Statistics (AISTATS)
2016
Cited alongside, same era.
M. Raginsky, A. Rakhlin, M. Tsao, Y. Wu, and A. Xu, “Information-theoretic analysis of stability and bias of learning algorithms,” in Proc. IEEE Information Theory Workshop (ITW)
A. Asadi, E. Abbe, and S. Verdu, “Chaining mutual information and tightening generalization bounds,” in Proc. Advances in Neural Information Processing Systems (NeurIPS)
2018
Later among the works it cites.
A. T. Lopez and V. Jog, “Generalization error bounds using Wasserstein distances,” in Proc. IEEE Information Theory Workshop (ITW)
2018
Later among the works it cites.
I. Issa and M. Gastpar, “Computable bounds on the exploration bias,” in Proc. IEEE Int. Symp. Information Theory (ISIT)
2018
Later among the works it cites.
A. Pensia, V. Jog, and P. Loh, “Generalization error bounds for noisy, iterative algorithms,” in Proc. IEEE Int. Symp. Information Theory (ISIT)
2018
Later among the works it cites.
W. Gao, S. Oh, and P. Viswanath, “Demystifying fixed k k -nearest neighbor information estimators,” IEEE Trans. Inform. Theory
2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals, “Understanding deep learning requires rethinking generalization,” in Proc. International Conference on Learning Representations (ICLR)
2017
Cited alongside, same era.
P. L. Bartlett, D. J. Foster, and M. J. Telgarsky, “Spectrally-normalized margin bounds for neural networks,” in Proc. Advances in Neural Information Processing Systems (NIPS)
2017
Cited alongside, same era.
A. Xu and M. Raginsky, “Information-theoretic analysis of generalization capability of learning algorithms,” in Proc. Advances in Neural Information Processing Systems (NIPS)
2017
Cited alongside, same era.
2017
Cited alongside, same era.
J. Jiao, Y. Han, and T. Weissman, “Dependence measures bounding the exploration bias for general measurements,” in Proc. IEEE Int. Symp. Information Theory (ISIT)
2017
Cited alongside, same era.
T. Young, D. Hazarika, S. Poria, and E. Cambria, “Recent trends in deep learning based natural language processing,” IEEE Computational Intelligence Magazine
2018
Cited alongside, same era.
R. Miotto, F. Wang, S. Wang, X. Jiang, and J. T. Dudley, “Deep learning for healthcare: review, opportunities and challenges,” Briefings in Bioinformatics
2018
Cited alongside, same era.
Later among the works it cites.
Y. Bu, S. Zou, and V. V. Veeravalli, “Tightening mutual information based bounds on generalization error,” in Proc. IEEE Int. Symp. Information Theory (ISIT)
2019
Closest in time.
2019
Closest in time.
H. Wang, M. Diaz, J. S. S. Filho, and F. P. Calmon, “An information-theoretic view of generalization via Wasserstein distance,” in Proc. IEEE Int. Symp. Information Theory (ISIT)
2019
Closest in time.
I. Issa, A. R. Esposito, and M. Gastpar, “Strengthened information-theoretic bounds on the generalization error,” in Proc. IEEE Int. Symp. Information Theory (ISIT)
2019
Closest in time.
J. Negrea, M. Haghifam, G. K. Dziugaite, A. Khisti, and D. M. Roy, “Information-theoretic generalization bounds for SGLD via data-dependent estimates,” in Proc. Advances in Neural Information Processing Systems (NeurIPS)
2019
Closest in time.
M. Rowland, J. Hron, Y. Tang, K. Choromanski, T. Sarlos, and A. Weller, “Orthogonal estimation of wasserstein distances,” in Proc. International Conference on Artifical Intelligence and Statistics (AISTATS)
2019
Closest in time.