Fetching the paper…
Reading the bibliography…
Machine learning models, while progressively advanced, rely heavily on the IID assumption, which is often unfulfilled in practice due to inevitable distribution shifts.
S. Wright, “Correlation and causation,” 1921
1921
Earlier work this paper cites.
V. Vapnik, “Principles of risk minimization for learning theory,” in NeurIPS 4 , vol. 4, 1991, pp. 831–838
1991
Earlier work this paper cites.
C.-Y. Chuang, A. Torralba, and S. Jegelka, “Estimating generalization under distribution shifts via domain-invariant representations,” in International Conference on Machine Learning . PMLR, 2020, pp. 1984–1994
1994
Earlier work this paper cites.
R. Kohavi et al. , “Scaling up the accuracy of naive-bayes classifiers: A decision-tree hybrid.” in Kdd , vol. 96, 1996, pp. 202–207
1996
Earlier work this paper cites.
M. Kukar, “Transductive reliability estimation for medical diagnosis,” Artificial Intelligence in Medicine , vol. 29, no. 1-2, pp. 81–106, 2003
2003
Earlier work this paper cites.
L. G. Neuberg, “Causality: models, reasoning, and inference, by judea pearl, cambridge university press, 2000,” Econometric Theory , vol. 19, no. 4, pp. 675–685, 2003
2003
Earlier work this paper cites.
O. Madani, D. Pennock, and G. Flake, “Co-validation: Using model disagreement on unlabeled data to validate classification algorithms,” Advances in neural information processing systems , vol. 17, 2004
2004
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in CVPR . Ieee, 2009, pp. 248–255
2009
Earlier work this paper cites.
S. Ben-David, J. Blitzer, K. Crammer, A. Kulesza, F. Pereira, and J. W. Vaughan, “A theory of learning from different domains,” Machine learning , vol. 79, no. 1-2, pp. 151–175, 2010
2010
Earlier work this paper cites.
A. Torralba and A. A. Efros, “Unbiased look at dataset bias,” in CVPR 2011 . IEEE, 2011, pp. 1521–1528
2011
Earlier work this paper cites.
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie, “The caltech-ucsd birds-200-2011 dataset,” 2011
2011
Earlier work this paper cites.
A. Maas, R. E. Daly, P. T. Pham, D. Huang, A. Y. Ng, and C. Potts, “Learning word vectors for sentiment analysis,” in Proceedings of the 49th annual meeting of the association for computational linguistics: Human language technologies , 2011, pp. 142–150
2011
Earlier work this paper cites.
A. Khosla, T. Zhou, T. Malisiewicz, A. A. Efros, and A. Torralba, “Undoing the damage of dataset bias,” in European Conference on Computer Vision . Springer, 2012, pp. 158–171
2012
Earlier work this paper cites.
R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Y. Ng, and C. Potts, “Recursive deep models for semantic compositionality over a sentiment treebank,” in Proceedings of the 2013 conference on empirical methods in natural language processing , 2013, pp. 1631–1642
2013
Earlier work this paper cites.
R. Shokri and V. Shmatikov, “Privacy-preserving deep learning,” in Proceedings of the 22nd ACM SIGSAC conference on computer and communications security , 2015, pp. 1310–1321
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
M. Long, Y. Cao, J. Wang, and M. Jordan, “Learning transferable features with deep adaptation networks,” in International conference on machine learning . PMLR, 2015, pp. 97–105
2015
Earlier work this paper cites.
Z. Liu, P. Luo, X. Wang, and X. Tang, “Deep learning face attributes in the wild,” in Proceedings of the IEEE international conference on computer vision , 2015, pp. 3730–3738
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
H. Namkoong and J. C. Duchi, “Stochastic gradient methods for distributionally robust optimization with f-divergences,” Advances in neural information processing systems , vol. 29, 2016
2016
Earlier work this paper cites.
G. Ros, L. Sellart, J. Materzynska, D. Vazquez, and A. M. Lopez, “The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 3234–3243
2016
Earlier work this paper cites.
S. R. Richter, V. Vineet, S. Roth, and V. Koltun, “Playing for data: Ground truth from computer games,” in Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part II 14 . Springer, 2016, pp. 102–118
2016
Earlier work this paper cites.
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for semantic urban scene understanding,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 3213–3223
2016
Earlier work this paper cites.
D. Anna Montoya, “House prices - advanced regression techniques,” 2016. [Online]. Available: https://kaggle.com/competitions/house-prices-advanced-regression-techniques
2016
Earlier work this paper cites.
D. Hendrycks and K. Gimpel, “A baseline for detecting misclassified and out-of-distribution examples in neural networks,” in International Conference on Learning Representations , 2016
2016
Earlier work this paper cites.
J. Peters, P. Bühlmann, and N. Meinshausen, “Causal inference by using invariant prediction: identification and confidence intervals,” Journal of the Royal Statistical Society. Series B (Statistical Methodology) , vol. 78, no. 5, pp. 947–1012, 2016. [Online]. Available: http://www.jstor.org/stable/44682904
2016
Earlier work this paper cites.
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang, “On large-batch training for deep learning: Generalization gap and sharp minima,” in International Conference on Learning Representations , 2016
2016
Earlier work this paper cites.
X. He, L. Liao, H. Zhang, L. Nie, X. Hu, and T.-S. Chua, “Neural collaborative filtering,” in Proceedings of the 26th international conference on world wide web , 2017, pp. 173–182
2017
Earlier work this paper cites.
D. Li, Y. Yang, Y.-Z. Song, and T. M. Hospedales, “Deeper, broader and artier domain generalization,” in ICCV , 2017, pp. 5542–5550
2017
Earlier work this paper cites.
M. Long, H. Zhu, J. Wang, and M. I. Jordan, “Deep transfer learning with joint adaptation networks,” in International conference on machine learning . PMLR, 2017, pp. 2208–2217
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
H. Venkateswara, J. Eusebio, S. Chakraborty, and S. Panchanathan, “Deep hashing network for unsupervised domain adaptation,” in CVPR , 2017, pp. 5018–5027
2017
Earlier work this paper cites.
M. Risdal, “New york city taxi trip duration,” 2017. [Online]. Available: https://kaggle.com/competitions/nyc-taxi-trip-duration
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
Q.-s. Zhang and S.-C. Zhu, “Visual interpretability for deep learning: a survey,” Frontiers of Information Technology & Electronic Engineering , vol. 19, no. 1, pp. 27–39, 2018
2018
Earlier work this paper cites.
N. Akhtar and A. Mian, “Threat of adversarial attacks on deep learning in computer vision: A survey,” Ieee Access , vol. 6, pp. 14 410–14 430, 2018
2018
Earlier work this paper cites.
A. Sinha, H. Namkoong, and J. Duchi, “Certifying some distributional robustness with principled adversarial training,” in International Conference on Learning Representations , 2018
2018
Earlier work this paper cites.
S. Beery, G. Van Horn, and P. Perona, “Recognition in terra incognita,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 456–473
2018
Earlier work this paper cites.
D. Hendrycks and T. Dietterich, “Benchmarking neural network robustness to common corruptions and perturbations,” in International Conference on Learning Representations , 2018
2018
Earlier work this paper cites.
A. Williams, N. Nangia, and S. Bowman, “A broad-coverage challenge corpus for sentence understanding through inference,” in Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) , 2018, pp. 1112–1122
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
N. Meinshausen, “Causality from a distributional robustness point of view,” in 2018 IEEE Data Science Workshop (DSW) . IEEE, 2018, pp. 6–10
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
J. C. Duchi, T. Hashimoto, and H. Namkoong, “Distributionally robust losses against mixture covariate shifts,” Under review , vol. 2, no. 1, 2019
2019
Earlier work this paper cites.
X. Peng, Q. Bai, X. Xia, Z. Huang, K. Saenko, and B. Wang, “Moment matching for multi-source domain adaptation,” in Proceedings of the IEEE/CVF international conference on computer vision , 2019, pp. 1406–1415
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
J. Irvin, P. Rajpurkar, M. Ko, Y. Yu, S. Ciurea-Ilcus, C. Chute, H. Marklund, B. Haghgoo, R. Ball, K. Shpanskaya et al. , “Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison,” in Proceedings of the AAAI conference on artificial intelligence , vol. 33, no. 01, 2019, pp. 590–597
2019
Earlier work this paper cites.
A. Barbu, D. Mayo, J. Alverio, W. Luo, C. Wang, D. Gutfreund, J. Tenenbaum, and B. Katz, “Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
B. Recht, R. Roelofs, L. Schmidt, and V. Shankar, “Do imagenet classifiers generalize to imagenet?” in International conference on machine learning . PMLR, 2019, pp. 5389–5400
2019
Earlier work this paper cites.
S. Sagawa, P. W. Koh, T. B. Hashimoto, and P. Liang, “Distributionally robust neural networks,” in International Conference on Learning Representations , 2019
2019
Cited alongside, same era.
P. Chaudhari, A. Choromanska, S. Soatto, Y. LeCun, C. Baldassi, C. Borgs, J. Chayes, L. Sagun, and R. Zecchina, “Entropy-sgd: Biasing gradient descent into wide valleys,” Journal of Statistical Mechanics: Theory and Experiment , vol. 2019, no. 12, p. 124018, 2019
2019
Cited alongside, same era.
R. Taori, A. Dave, V. Shankar, N. Carlini, B. Recht, and L. Schmidt, “When robustness doesn’t promote robustness: Synthetic vs. natural distribution shifts on imagenet,” 2019
2019
Cited alongside, same era.
I. Gulrajani and D. Lopez-Paz, “In search of lost domain generalization,” in International Conference on Learning Representations , 2020
2020
Cited alongside, same era.
J. Lee, C. Liu, J. Kim, Z. Chen, Y. Sun, J. R. Rogers, W. K. Chung, and C. Weng, “Deep learning for rare disease: A scoping review,” Journal of Biomedical Informatics , p. 104227, 2022
2022
Later among the works it cites.
C. Baek, Y. Jiang, A. Raghunathan, and J. Z. Kolter, “Agreement-on-the-line: Predicting the performance of neural networks under distribution shift,” Advances in Neural Information Processing Systems , vol. 35, pp. 19 274–19 289, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
D. Hendrycks, S. Basart, M. Mazeika, A. Zou, J. Kwon, M. Mostajabi, J. Steinhardt, and D. Song, “Scaling out-of-distribution detection for real-world settings,” in International Conference on Machine Learning . PMLR, 2022, pp. 8759–8773
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Shen, P. Cui, T. Zhang, and K. Kunag, “Stable learning via sample reweighting,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 04, 2020, pp. 5692–5699
2020
Cited alongside, same era.
K. Kuang, R. Xiong, P. Cui, S. Athey, and B. Li, “Stable prediction with model misspecification and agnostic distribution shift,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 04, 2020, pp. 4485–4492
2020
Cited alongside, same era.
D. C. Castro, I. Walker, and B. Glocker, “Causality matters in medical imaging,” Nature Communications , vol. 11, no. 1, p. 3673, 2020
2020
Cited alongside, same era.
P. Foret, A. Kleiner, H. Mobahi, and B. Neyshabur, “Sharpness-aware minimization for efficiently improving generalization,” in International Conference on Learning Representations , 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
D. Hendrycks, X. Liu, E. Wallace, A. Dziedzic, R. Krishnan, and D. Song, “Pretrained transformers improve out-of-distribution robustness,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , 2020, pp. 2744–2751
2020
Cited alongside, same era.
Z. Shen, P. Cui, J. Liu, T. Zhang, B. Li, and Z. Chen, “Stable learning via differentiated variable decorrelation,” in Proceedings of the 26th acm sigkdd international conference on knowledge discovery & data mining , 2020, pp. 2185–2193
2020
Cited alongside, same era.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial networks,” Communications of the ACM , vol. 63, no. 11, pp. 139–144, 2020
2020
Cited alongside, same era.
2022
Later among the works it cites.
S. Jesus, J. Pombal, D. Alves, A. Cruz, P. Saleiro, R. Ribeiro, J. Gama, and P. Bizarro, “Turning the tables: Biased, imbalanced, dynamic tabular datasets for ml evaluation,” Advances in Neural Information Processing Systems , vol. 35, pp. 33 563–33 575, 2022
2022
Later among the works it cites.
B. Y. Idrissi, M. Arjovsky, M. Pezeshki, and D. Lopez-Paz, “Simple data balancing achieves competitive worst-group-accuracy,” in Conference on Causal Learning and Reasoning . PMLR, 2022, pp. 336–351
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
Y. Yu, Z. Yang, A. Wei, Y. Ma, and J. Steinhardt, “Predicting out-of-distribution error with the projection norm,” in International Conference on Machine Learning . PMLR, 2022, pp. 25 721–25 746
2022
Later among the works it cites.
A. Kirsch and Y. Gal, “A note on” assessing generalization of sgd via disagreement”,” Transactions on Machine Learning Research , 2022
2022
Later among the works it cites.
L. Chen, M. Zaharia, and J. Y. Zou, “Is unsupervised performance estimation impossible when both covariates and labels shift?” in NeurIPS 2022 Workshop on Distribution Shifts: Connecting Methods and Applications , 2022
2022
Later among the works it cites.
N. Thams, M. Oberst, and D. Sontag, “Evaluating robustness to dataset shift via parametric robustness sets,” Advances in Neural Information Processing Systems , vol. 35, pp. 16 877–16 889, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
J. Liu, J. Wu, R. Pi, R. Xu, X. Zhang, B. Li, and P. Cui, “Measure the predictive heterogeneity,” in The Eleventh International Conference on Learning Representations , 2022
2022
Later among the works it cites.
D. Arpit, H. Wang, Y. Zhou, and C. Xiong, “Ensemble of averages: Improving model selection and boosting performance in domain generalization,” Advances in Neural Information Processing Systems , vol. 35, pp. 8265–8277, 2022
2022
Later among the works it cites.
X. Chu, Y. Jin, W. Zhu, Y. Wang, X. Wang, S. Zhang, and H. Mei, “Dna: Domain generalization with diversified neural averaging,” in International Conference on Machine Learning . PMLR, 2022, pp. 4010–4034
2022
Later among the works it cites.
M. Wortsman, G. Ilharco, S. Y. Gadre, R. Roelofs, R. Gontijo-Lopes, A. S. Morcos, H. Namkoong, A. Farhadi, Y. Carmon, S. Kornblith et al. , “Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time,” in International Conference on Machine Learning . PMLR, 2022, pp. 23 965–23 998
2022
Later among the works it cites.
A. Rame, M. Kirchmeyer, T. Rahier, A. Rakotomamonjy, P. Gallinari, and M. Cord, “Diverse weight averaging for out-of-distribution generalization,” Advances in Neural Information Processing Systems , vol. 35, pp. 10 821–10 836, 2022
2022
Later among the works it cites.
Z. Li, K. Ren, X. Jiang, Y. Shen, H. Zhang, and D. Li, “Simple: Specialized model-sample matching for domain generalization,” in The Eleventh International Conference on Learning Representations , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
X. Zhang, Y. He, T. Wang, J. Qi, H. Yu, Z. Wang, J. Peng, R. Xu, Z. Shen, Y. Niu et al. , “Nico challenge: Out-of-distribution generalization for image recognition challenges,” in European Conference on Computer Vision . Springer, 2022, pp. 433–450
2022
Later among the works it cites.
A. J. Andreassen, Y. Bahri, B. Neyshabur, and R. Roelofs, “The evolution of out-of-distribution robustness throughout fine-tuning,” Transactions on Machine Learning Research , 2022
2022
Later among the works it cites.
2023
Later among the works it cites.
H. Yu, P. Cui, Y. He, Z. Shen, Y. Lin, R. Xu, and X. Zhang, “Stable learning via sparse variable independence,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 37, no. 9, 2023, pp. 10 998–11 006
2023
Later among the works it cites.
X. Zhang, Y. He, R. Xu, H. Yu, Z. Shen, and P. Cui, “Nico++: Towards better benchmarking for domain generalization,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 16 036–16 047
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Kulinski and D. I. Inouye, “Towards explaining distribution shifts,” in International Conference on Machine Learning . PMLR, 2023, pp. 17 931–17 952
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
P. Trivedi, D. Koutra, and J. J. Thiagarajan, “A closer look at scoring functions and generalization prediction,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
X. Zhang, R. Xu, H. Yu, H. Zou, and P. Cui, “Gradient norm aware minimization seeks first-order flatness and improves generalization,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 20 247–20 257
2023
Later among the works it cites.
A. Rame, K. Ahuja, J. Zhang, M. Cord, L. Bottou, and D. Lopez-Paz, “Model ratatouille: Recycling diverse models for out-of-distribution generalization,” 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.