Fetching the paper…
Reading the bibliography…
We derive upper bounds on the generalization error of learning algorithms based on their \emph{algorithmic transport cost}: the expected Wasserstein distance between the output hypothesis and the output hypothesis conditioned on an input example.
Théorie asymptotique de la décision statistique
Le Cam, L. M. (1969) · 1969
Earlier work this paper cites.
An overview of statistical learning theory
Vapnik, V. N. (1999) · 1999
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Bartlett, P. L. and Mendelson, S. (2002) · 2002
Earlier work this paper cites.
Stability and generalization
Bousquet, O. and Elisseeff, A. (2002) · 2002
Earlier work this paper cites.
On choosing and bounding probability metrics
Gibbs, A. L. and Su, F. E. (2002) · 2002
Earlier work this paper cites.
Algorithmic luckiness
Herbrich, R. and Williamson, R. C. (2002) · 2002
Earlier work this paper cites.
Covering number bounds of certain regularized linear function classes
Zhang, T. (2002) · 2002
Earlier work this paper cites.
The covering number in learning theory
Zhou, D.-X. (2002) · 2002
Earlier work this paper cites.
Stochastic approximation and recursive algorithms and applications
Kushner, H. and Yin, G. G. (2003) · 2003
Earlier work this paper cites.
Local rademacher complexities
Bartlett, P. L., Bousquet, O., Mendelson, S., et al. (2005) · 2005
Earlier work this paper cites.
Tutorial on practical prediction theory for classification
Langford, J. (2005) · 2005
Earlier work this paper cites.
Tighter pac-bayes bounds
Ambroladze, A., Parrado-Hernández, E., and Shawe-taylor, J. S. (2007) · 2007
Earlier work this paper cites.
Optimal transport: old and new
Villani, C. (2008) · 2008
Earlier work this paper cites.
Learnability, stability and uniform convergence
Shalev-Shwartz, S., Shamir, O., Srebro, N., and Sridharan, K. (2010) · 2010
Earlier work this paper cites.
Elements of information theory
Cover, T. M. and Thomas, J. A. (2012) · 2012
Cited alongside, same era.
Foundations of machine learning
Mohri, M., Rostamizadeh, A., and Talwalkar, A. (2012) · 2012
Cited alongside, same era.
Robustness and generalization
Xu, H. and Mannor, S. (2012) · 2012
Cited alongside, same era.
The nature of statistical learning theory
Vapnik, V. (2013) · 2013
Cited alongside, same era.
Algorithmic stability and uniform generalization
Alabdulmohsin, I. M. (2015) · 2015
Cited alongside, same era.
Generalization in adaptive data analysis and holdout reuse
Dwork, C., Feldman, V., Hardt, M., Pitassi, T., Reingold, O., and Roth, A. (2015) · 2015
Cited alongside, same era.
Algorithmic stability for adaptive data analysis
An Information-Theoretic Route from Generalization in Expectation to Generalization in Probability
Alabdulmohsin, I. (2017) · 2017
Later among the works it cites.
Wasserstein generative adversarial networks
Arjovsky, M., Chintala, S., and Bottou, L. (2017) · 2017
Later among the works it cites.
Optimal transport for domain adaptation
Courty, N., Flamary, R., Tuia, D., and Rakotomamonjy, A. (2017) · 2017
Later among the works it cites.
Improved training of wasserstein gans
Gulrajani, I., Ahmed, F., Arjovsky, M., Dumoulin, V., and Courville, A. C. (2017) · 2017
Later among the works it cites.
Minimax statistical learning and domain adaptation with wasserstein distances
Lee, J. and Raginsky, M. (2017) · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bassily, R., Nissim, K., Smith, A., Steinke, T., Stemmer, U., and Ullman, J. (2016) · 2016
Cited alongside, same era.
Distributionally robust stochastic optimization with wasserstein distance
Gao, R. and Kleywegt, A. J. (2016) · 2016
Cited alongside, same era.
Information-theoretic analysis of stability and bias of learning algorithms
Raginsky, M., Rakhlin, A., Tsao, M., Wu, Y., and Xu, A. (2016) · 2016
Cited alongside, same era.
Controlling bias in adaptive data analysis using information theory
Russo, D. and Zou, J. (2016) · 2016
Cited alongside, same era.
f f -divergence inequalities
Sason, I. and Verdú, S. (2016) · 2016
Cited alongside, same era.
On-average kl-privacy and its equivalence to generalization for max-entropy mechanisms
Wang, Y.-X., Lei, J., and Fienberg, S. E. (2016) · 2016
Cited alongside, same era.
Liu, T., Lugosi, G., Neu, G., and Tao, D. (2017) · 2017
Later among the works it cites.
Computational optimal transport
Peyré, G., Cuturi, M., et al. (2017) · 2017
Later among the works it cites.
Strong data-processing inequalities for channels and bayesian networks
Polyanskiy, Y. and Wu, Y. (2017) · 2017
Later among the works it cites.
Information-theoretic analysis of generalization capability of learning algorithms
Xu, A. and Raginsky, M. (2017) · 2017
Later among the works it cites.
Information theoretic guarantees for empirical risk minimization with applications to model selection and large-scale optimization
Alabdulmohsin, I. (2018) · 2018
Closest in time.
Learners that use little information
Bassily, R., Moran, S., Nachum, I., Shafer, J., and Yehudayoff, A. (2018) · 2018
Closest in time.
A direct sum result for the information complexity of learning
Nachum, I., Shafer, J., and Yehudayoff, A. (2018) · 2018
Closest in time.
Pac-bayes bounds for stable algorithms with instance-dependent priors
Rivasplata, O., Parrado-Hernandez, E., Shaws-Taylor, J., Sun, S., and Szepesvari, C. (2018) · 2018
Closest in time.