Fetching the paper…
Reading the bibliography…
Transfer learning, or domain adaptation, is concerned with machine learning problems in which training and testing data come from possibly different probability distributions.
V. N. Vapnik and A. Y. Chervonenkis, “On the uniform convergence of relative frequencies of events to their probabilities,” Theory of Probability and Its Applications , 1971
1971
Earlier work this paper cites.
M. D. Donsker and S. S. Varadhan, “Asymptotic evaluation of certain markov process expectations for large time, i,” Communications on Pure and Applied Mathematics , vol. 28, no. 1, pp. 1–47, 1975
1975
Earlier work this paper cites.
L. Devroye and T. Wagner, “Distribution-free inequalities for the deleted and holdout error estimates,” IEEE Transactions on Information Theory , vol. 25, no. 2, pp. 202–207, 1979
1979
Earlier work this paper cites.
R. Moddemeijer, “On estimation of entropy and mutual information of continuous distributions,” Signal processing , vol. 16, no. 3, pp. 233–248, 1989
1989
Earlier work this paper cites.
P. J. Huber, “Robust estimation of a location parameter,” in Breakthroughs in statistics: Methodology and distribution . Springer, 1992, pp. 492–518
1992
Earlier work this paper cites.
M. Kearns and D. Ron, “Algorithmic stability and sanity-check bounds for leave-one-out cross-validation,” in Proceedings of the tenth annual conference on Computational learning theory , 1997, pp. 152–162
1997
Earlier work this paper cites.
Y. Freund and R. E. Schapire, “A decision-theoretic generalization of online learning and an application to boosting,” Journal of computer and system sciences , vol. 55, no. 1, pp. 119–139, 1997
1997
Earlier work this paper cites.
D. A. McAllester, “Some pac-bayesian theorems,” Machine Learning , vol. 37, no. 3, pp. 355–363, 1999
1999
Earlier work this paper cites.
V. Vapnik, The nature of statistical learning theory . Springer science & business media, 1999
1999
Earlier work this paper cites.
S. G. Bobkov and F. Götze, “Exponential integrability and transportation cost related to logarithmic sobolev inequalities,” Journal of Functional Analysis , vol. 163, no. 1, pp. 1–28, 1999
1999
Earlier work this paper cites.
O. Bousquet and A. Elisseeff, “Stability and generalization,” Journal of Machine Learning Research , vol. 2, no. Mar, pp. 499–526, 2002
2002
Earlier work this paper cites.
S. P. Boyd and L. Vandenberghe, Convex optimization . Cambridge University Press, 2004
2004
Earlier work this paper cites.
P. L. Bartlett and S. Mendelson, “Empirical minimization,” Probability theory and related fields , vol. 135, no. 3, pp. 311–334, 2006
2006
Earlier work this paper cites.
P. L. Bartlett, M. I. Jordan, and J. D. McAuliffe, “Convexity, classification, and risk bounds,” Journal of the American Statistical Association , vol. 101, no. 473, pp. 138–156, 2006
2006
Earlier work this paper cites.
J. Huang, A. Gretton, K. Borgwardt, B. Schölkopf, and A. Smola, “Correcting sample selection bias by unlabeled data,” Advances in neural information processing systems , vol. 19, 2006
2006
Earlier work this paper cites.
J. Blitzer, K. Crammer, A. Kulesza, F. Pereira, and J. Wortman, “Learning bounds for domain adaptation,” Advances in neural information processing systems , vol. 20, 2007
2007
Earlier work this paper cites.
W. Dai, Q. Yang, G.-R. Xue, and Y. Yu, “Boosting for transfer learning,” in Proceedings of the 24th international conference on Machine learning . ACM, 2007, pp. 193–200
2007
Earlier work this paper cites.
2009
Earlier work this paper cites.
A. Gretton, A. Smola, J. Huang, M. Schmittfull, K. Borgwardt, and B. Schölkopf, “Covariate shift by kernel mean matching,” Dataset shift in machine learning , vol. 3, no. 4, p. 5, 2009
2009
Earlier work this paper cites.
S. Ben-David, J. Blitzer, K. Crammer, A. Kulesza, F. Pereira, and J. W. Vaughan, “A theory of learning from different domains,” Machine Learning , vol. 79, no. 1, pp. 151–175, May 2010
2010
Earlier work this paper cites.
R. M. Dudley, “Distances of probability measures and random variables,” in Selected Works of RM Dudley . Springer, 2010, pp. 28–37
2010
Earlier work this paper cites.
E. Eaton et al. , “Selective transfer between learning tasks using task-based boosting,” in Twenty-Fifth AAAI Conference on Artificial Intelligence , 2011
2011
Earlier work this paper cites.
E. Moulines and F. Bach, “Non-asymptotic analysis of stochastic approximation algorithms for machine learning,” Advances in neural information processing systems , vol. 24, 2011
2011
Earlier work this paper cites.
M. Welling and Y. W. Teh, “Bayesian learning via stochastic gradient langevin dynamics,” in Proceedings of the 28th international conference on machine learning (ICML-11) , 2011, pp. 681–688
2011
Earlier work this paper cites.
M. Schmidt, N. L. Roux, and F. R. Bach, “Convergence rates of inexact proximal-gradient methods for convex optimization,” in Advances in neural information processing systems , 2011, pp. 1458–1466
2011
Earlier work this paper cites.
2011
Earlier work this paper cites.
B. Recht, C. Re, S. Wright, and F. Niu, “Hogwild: A lock-free approach to parallelizing stochastic gradient descent,” Advances in neural information processing systems , vol. 24, 2011
2011
Earlier work this paper cites.
H. Xu and S. Mannor, “Robustness and generalization,” Machine learning , vol. 86, no. 3, pp. 391–423, 2012
2012
Earlier work this paper cites.
C. Zhang, L. Zhang, and J. Ye, “Generalization bounds for domain adaptation,” in Advances in neural information processing systems , 2012, pp. 3320–3328
2012
Earlier work this paper cites.
B. K. Sriperumbudur, K. Fukumizu, A. Gretton, B. Schölkopf, G. R. Lanckriet et al. , “On the empirical estimation of integral probability metrics,” Electronic Journal of Statistics , vol. 6, pp. 1550–1599, 2012
2012
Earlier work this paper cites.
M. Sugiyama, T. Suzuki, and T. Kanamori, Density ratio estimation in machine learning . Cambridge University Press, 2012
2012
Cited alongside, same era.
M. Long, J. Wang, G. Ding, S. J. Pan, and S. Y. Philip, “Adaptation regularization: A general framework for transfer learning,” IEEE Transactions on Knowledge and Data Engineering , vol. 26, no. 5, pp. 1076–1089, 2013
2013
Cited alongside, same era.
I. Kuzborskij and F. Orabona, “Stability and hypothesis transfer learning,” in International Conference on Machine Learning . PMLR, 2013, pp. 942–950
2013
Cited alongside, same era.
P. Germain, A. Habrard, F. Laviolette, and E. Morvant, “A pac-bayesian approach for domain adaptation with specialization to linear classifiers,” in International conference on machine learning . PMLR, 2013, pp. 738–746
2013
Cited alongside, same era.
H. Wang, M. Diaz, J. C. S. Santos Filho, and F. P. Calmon, “An information-theoretic view of generalization via wasserstein distance,” in 2019 IEEE International Symposium on Information Theory (ISIT) . IEEE, 2019, pp. 577–581
2019
Later among the works it cites.
J. Negrea, M. Haghifam, G. K. Dziugaite, A. Khisti, and D. M. Roy, “Information-theoretic generalization bounds for sgld via data-dependent estimates,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Later among the works it cites.
J. Van Baar, A. Sullivan, R. Cordorel, D. Jha, D. Romeres, and D. Nikovski, “Sim-to-real transfer learning using robustified controllers in robotic tasks involving complex dynamics,” in 2019 International Conference on Robotics and Automation (ICRA) . IEEE, 2019, pp. 6001–6007
2019
Later among the works it cites.
2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Raginsky, I. Sason et al. , “Concentration of measure inequalities in information theory, communications, and coding,” Foundations and Trends® in Communications and Information Theory , vol. 10, no. 1-2, pp. 1–246, 2013
2013
Cited alongside, same era.
S. Boucheron, G. Lugosi, and P. Massart, Concentration inequalities: A nonasymptotic theory of independence . OUP Oxford, Feb. 2013
2013
Cited alongside, same era.
S. Shalev-Shwartz and S. Ben-David, Understanding machine learning: From theory to algorithms . Cambridge University Press, 2014
2014
Cited alongside, same era.
C. Dwork, V. Feldman, M. Hardt, T. Pitassi, O. Reingold, and A. L. Roth, “Preserving statistical validity in adaptive data analysis,” in Proceedings of the forty-seventh annual ACM symposium on Theory of computing , 2015, pp. 117–126
2015
Cited alongside, same era.
T. Van Erven, P. D. Grünwald, N. A. Mehta, M. D. Reid, and R. C. Williamson, “Fast rates in statistical and online learning,” The Journal of Machine Learning Research , vol. 16, no. 1, pp. 1793–1861, 2015
2015
Cited alongside, same era.
D. Russo and J. Zou, “Controlling bias in adaptive data analysis using information theory,” in Artificial Intelligence and Statistics . PMLR, 2016, pp. 1232–1240
2016
Cited alongside, same era.
M. Raginsky, A. Rakhlin, M. Tsao, Y. Wu, and A. Xu, “Information-theoretic analysis of stability and bias of learning algorithms,” in 2016 IEEE Information Theory Workshop (ITW) . IEEE, 2016, pp. 26–30
2016
Cited alongside, same era.
2016
Cited alongside, same era.
Later among the works it cites.
F. Salehi, E. Abbasi, and B. Hassibi, “The impact of regularization on high-dimensional logistic regression,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Later among the works it cites.
S. Hanneke and S. Kpotufe, “On the value of target data in transfer learning,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Later among the works it cites.
B. Wang, J. Mendez, M. Cai, and E. Eaton, “Transfer learning via minimizing the performance gap between domains,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Later among the works it cites.
X. Wu, J. H. Manton, U. Aickelin, and J. Zhu, “Information-theoretic analysis for transfer learning,” in 2020 IEEE International Symposium on Information Theory (ISIT) . IEEE, 2020, pp. 2819–2824
2020
Later among the works it cites.
T. Steinke and L. Zakynthinou, “Reasoning about generalization via conditional mutual information,” in Conference on Learning Theory . PMLR, 2020, pp. 3437–3452
2020
Later among the works it cites.
M. Haghifam, J. Negrea, A. Khisti, D. M. Roy, and G. K. Dziugaite, “Sharpened generalization bounds based on conditional mutual information and an application to noisy, iterative algorithms,” Advances in Neural Information Processing Systems , vol. 33, pp. 9925–9935, 2020
2020
Later among the works it cites.
P. D. Grünwald and N. A. Mehta, “Fast rates for general unbounded loss functions: From erm to generalized bayes.” J. Mach. Learn. Res. , vol. 21, pp. 56–1, 2020
2020
Later among the works it cites.
J. Zhu, “Semi-supervised learning: The case when unlabeled data is equally useful,” in Conference on Uncertainty in Artificial Intelligence . PMLR, 2020, pp. 709–718
2020
Later among the works it cites.
B. Rodríguez Gálvez, G. Bassi, R. Thobaben, and M. Skoglund, “Tighter expected generalization error bounds via wasserstein distance,” Advances in Neural Information Processing Systems , vol. 34, pp. 19 109–19 121, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
M. A. Morid, A. Borjali, and G. Del Fiol, “A scoping review of transfer learning research on medical image analysis using imagenet,” Computers in biology and medicine , vol. 128, p. 104115, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
B. Liu, Y. Cai, Y. Guo, and X. Chen, “Transtailor: Pruning the pre-trained model for improved transfer learning,” in Proceedings of the AAAI conference on artificial intelligence , vol. 35, no. 10, 2021, pp. 8627–8634
2021
Later among the works it cites.
K. You, Y. Liu, J. Wang, and M. Long, “Logme: Practical assessment of pre-trained models for transfer learning,” in International Conference on Machine Learning . PMLR, 2021, pp. 12 133–12 143
2021
Later among the works it cites.
S. Niu, M. Liu, Y. Liu, J. Wang, and H. Song, “Distant domain transfer learning for medical imaging,” IEEE Journal of Biomedical and Health Informatics , vol. 25, no. 10, pp. 3784–3793, 2021
2021
Later among the works it cites.
Y. Bu, G. Aminian, L. Toni, G. W. Wornell, and M. Rodrigues, “Characterizing and understanding the generalization error of transfer learning with gibbs algorithm,” in International Conference on Artificial Intelligence and Statistics . PMLR, 2022, pp. 8673–8699
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
G. Aminian, M. Abroshan, M. M. Khalili, L. Toni, and M. Rodrigues, “An information-theoretical approach to semi-supervised learning under covariate-shift,” in International Conference on Artificial Intelligence and Statistics . PMLR, 2022, pp. 7433–7449
2022
Closest in time.
A. R. Esposito and M. Gastpar, “From generalisation error to transportation-cost inequalities and back,” in 2022 IEEE International Symposium on Information Theory (ISIT) . IEEE, 2022, pp. 294–299
2022
Closest in time.
M. Haghifam, B. Rodríguez-Gálvez, R. Thobaben, M. Skoglund, D. M. Roy, and G. K. Dziugaite, “Limitations of information-theoretic generalization bounds for gradient descent methods in stochastic convex optimization,” in International Conference on Algorithmic Learning Theory . PMLR, 2023, pp. 663–706
2023
Closest in time.
A. Dembo, T. M. Cover, and J. A. Thomas, “Information theoretic inequalities,” IEEE Transactions on Information Theory , vol. 37, no. 6, pp. 1501–1518, 1991
2023
Closest in time.
B. Gong, Y. Shi, F. Sha, and K. Grauman, “Geodesic flow kernel for unsupervised domain adaptation,” in 2012 IEEE conference on computer vision and pattern recognition . IEEE, 2012, pp. 2066–2073
2073
Closest in time.