Fetching the paper…
Reading the bibliography…
To avoid failures on out-of-distribution data, recent works have sought to extract features that have an invariant or stable relationship with the label across domains, discarding "spurious" or unstable features whose relationship with the label changes across domains.
Arjovsky, M., Bottou, L., Gulrajani, I., and Lopez-Paz, D. (2020) · 1907
Earlier work this paper cites.
Learning from noisy labels with distillation
Li, Y., Yang, J., Song, Y., Cao, L., Luo, J., and Li, L.-J. (2017b) · 1918
Earlier work this paper cites.
Masstabinvariante korrelationstheorie
Hoeffding, W. (1940) · 1940
Earlier work this paper cites.
Masstabinvariante korrelationsmasse für diskontinuierliche verteilungen
Hoeffding, W. (1941) · 1941
Earlier work this paper cites.
Sur les tableaux de corrélation dont les marges sont données
Fréchet, M. (1951) · 1951
Earlier work this paper cites.
The strength of weak learnability
Schapire, R. E. (1990) · 1990
Earlier work this paper cites.
Principles of risk minimization for learning theory
Vapnik, V. (1991) · 1991
Earlier work this paper cites.
Combining labeled and unlabeled data with co-training
Blum, A. and Mitchell, T. (1998) · 1998
Earlier work this paper cites.
Statistical Learning Theory
Vapnik, V. N. (1998) · 1998
Earlier work this paper cites.
Multi-relational learning, text mining, and semi-supervised learning for functional genomics
Krogel, M.-A. and Scheffer, T. (2004) · 2004
Earlier work this paper cites.
Measuring statistical dependence with Hilbert-Schmidt norms
Gretton, A., Bousquet, O., Smola, A., and Schölkopf, B. (2005) · 2005
Earlier work this paper cites.
Model compression
Buciluǎ, C., Caruana, R., and Niculescu-Mizil, A. (2006) · 2006
Earlier work this paper cites.
One-shot learning of object categories
Fei-Fei, L., Fergus, R., and Perona, P. (2006) · 2006
Earlier work this paper cites.
Learning bounds for domain adaptation
Blitzer, J., Crammer, K., Kulesza, A., Pereira, F., and Wortman, J. (2007) · 2007
Earlier work this paper cites.
Empirical comparison of “hard” and “soft” label propagation for relational classification
Galstyan, A. and Cohen, P. R. (2008) · 2007
Earlier work this paper cites.
In search of lost domain generalization
Gulrajani, I. and Lopez-Paz, D. (2020) · 2007
Earlier work this paper cites.
Covariate shift adaptation by importance weighted cross validation
Sugiyama, M., Krauledat, M., and Müller, K.-R. (2007) · 2007
Earlier work this paper cites.
Domain adaptation with multiple sources
Mansour, Y., Mohri, M., and Rostamizadeh, A. (2008) · 2008
Earlier work this paper cites.
Discriminative learning under covariate shift
Bickel, S., Brückner, M., and Scheffer, T. (2009) · 2009
Earlier work this paper cites.
Covariate shift by kernel mean matching
Gretton, A., Smola, A., Huang, J., Schmittfull, M., Borgwardt, K., and Schölkopf, B. (2009) · 2009
Earlier work this paper cites.
A theory of learning from different domains
Ben-David, S., Blitzer, J., Crammer, K., Kulesza, A., Pereira, F., and Vaughan, J. W. (2010) · 2010
Earlier work this paper cites.
Generalizing from several related classification tasks to a new unlabeled sample
Blanchard, G., Lee, G., and Scott, C. (2011) · 2011
Earlier work this paper cites.
Machine learning in non-stationary environments: Introduction to covariate shift adaptation
Sugiyama, M. and Kawanabe, M. (2012) · 2012
Earlier work this paper cites.
Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks
Lee, D.-H. et al. (2013) · 2013
Earlier work this paper cites.
Domain generalization via invariant feature representation
Muandet, K., Balduzzi, D., and Schölkopf, B. (2013) · 2013
Cited alongside, same era.
Learning with noisy labels
Natarajan, N., Dhillon, I. S., Ravikumar, P. K., and Tewari, A. (2013) · 2013
Cited alongside, same era.
Classification with asymmetric label noise: Consistency and maximal denoising
Scott, C., Blanchard, G., and Handy, G. (2013) · 2013
Cited alongside, same era.
Do deep nets really need to be deep?
Ba, J. and Caruana, R. (2014) · 2014
Cited alongside, same era.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J. (2015) · 2015
Cited alongside, same era.
Classification with asymmetric label noise: Consistency and maximal denoising
Blanchard, G., Flaska, M., Handy, G., Pozzi, S., and Scott, C. (2016) · 2016
Exploiting domain-specific features to enhance domain generalization
Bui, M.-H., Tran, T., Tran, A., and Phung, D. (2021) · 2021
Later among the works it cites.
Unit-level surprise in neural networks
Eastwood, C., Mason, I., and Williams, C. (2021) · 2021
Later among the works it cites.
Test-time classifier adjustment module for model-agnostic domain generalization
Iwasawa, Y. and Matsuo, Y. (2021) · 2021
Later among the works it cites.
WILDS: A benchmark of in-the-wild distribution shifts
Koh, P. W., Sagawa, S., Marklund, H., Xie, S. M., Zhang, M., Balsubramani, A., Hu, W., Yasunaga, M., Phillips, R. L., Gao, I., Lee, T., David, E., Stavness, I., Guo, W., Earnshaw, B. A., Haque, I. S., Beery, S., Leskovec, J., Kundaje, A., Pierson, E., Levine, S., Finn, C., and Liang, P. (2021) · 2021
Later among the works it cites.
Out-of-distribution generalization via risk extrapolation (REx)
Krueger, D., Caballero, E., Jacobsen, J.-H., Zhang, A., Binas, J., Zhang, D., Priol, R. L., and Courville, A. (2021) · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Classifier calibration
Flach, P. A. (2016) · 2016
Cited alongside, same era.
Causal inference by using invariant prediction: identification and confidence intervals
Peters, J., Bühlmann, P., and Meinshausen, N. (2016) · 2016
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S. (2017) · 2017
Cited alongside, same era.
On calibration of modern neural networks
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. (2017) · 2017
Cited alongside, same era.
From detection of individual metastases to classification of lymph node status at the patient level: the camelyon17 challenge
Bandi, P., Geessink, O., Manson, Q., Van Dijk, M., Balkenhol, M., Hermsen, M., Bejnordi, B. E., Lee, B., Paeng, K., Zhong, A., et al. (2018) · 2018
Cited alongside, same era.
Recognition in terra incognita
Beery, S., Van Horn, G., and Perona, P. (2018) · 2018
Cited alongside, same era.
Understanding the failure modes of out-of-distribution generalization
Nagarajan, V., Andreassen, A., and Neyshabur, B. (2021) · 2021
Later among the works it cites.
Gradient starvation: A learning proclivity in neural networks
Pezeshki, M., Kaba, O., Bengio, Y., Courville, A. C., Precup, D., and Lajoie, G. (2021) · 2021
Later among the works it cites.
Anchor regression: Heterogeneous data meet causality
Rothenhäusler, D., Meinshausen, N., Bühlmann, P., and Peters, J. (2021) · 2021
Later among the works it cites.
Counterfactual invariance to spurious correlations: Why and how to pass stress tests
Veitch, V., D’Amour, A., Yadlowsky, S., and Eisenstein, J. (2021) · 2021
Later among the works it cites.
On calibration and out-of-domain generalization
Wald, Y., Feder, A., Greenfeld, D., and Shalit, U. (2021) · 2021
Later among the works it cites.
Invariant and transportable representations for anti-causal domain shifts
Jiang, Y. and Veitch, V. (2022) · 2022
Later among the works it cites.
Last layer re-training is sufficient for robustness to spurious correlations
Kirichenko, P., Izmailov, P., and Wilson, A. G. (2022) · 2022
Later among the works it cites.
Causally motivated shortcut removal using auxiliary labels
Makar, M., Packer, B., Moldovan, D., Blalock, D., Halpern, Y., and D’Amour, A. (2022) · 2022
Later among the works it cites.
Out-of-distribution generalization in the presence of nuisance-induced spurious correlations
Puli, A. M., Zhang, L. H., Oermann, E. K., and Ranganath, R. (2022) · 2022
Later among the works it cites.
Fishr: Invariant gradient variances for out-of-distribution generalization
Rame, A., Dancette, C., and Cord, M. (2022) · 2022
Later among the works it cites.
Rosenfeld, E., Ravikumar, P., and Risteski, A. (2022) · 2022
Later among the works it cites.
If your data distribution shifts, use self-learning
Rusak, E., Schneider, S., Pachitariu, G., Eck, L., Gehler, P. V., Bringmann, O., Brendel, W., and Bethge, M. (2022) · 2022
Later among the works it cites.
Causality for machine learning
Schölkopf, B. (2022) · 2022
Later among the works it cites.
Learning from noisy labels with deep neural networks: A survey
Song, H., Kim, M., Park, D., Shin, Y., and Lee, J.-G. (2022) · 2022
Later among the works it cites.
Beyond invariance: Test-time label-shift adaptation for distributions with" spurious" correlations
Sun, Q., Murphy, K., Ebrahimi, S., and D’Amour, A. (2022) · 2022
Later among the works it cites.
Rich feature construction for the optimization-generalization dilemma
Zhang, J., Lopez-Paz, D., and Bottou, L. (2022) · 2022
Later among the works it cites.
Causally motivated multi-shortcut identification & removal
Zheng, J. and Makar, M. (2022) · 2022
Later among the works it cites.
Classifier calibration: a survey on how to assess and improve predicted class probabilities
Silva Filho, T., Song, H., Perello-Nieto, M., Santos-Rodriguez, R., Kull, M., and Flach, P. (2023) · 2023
Closest in time.