Fetching the paper…
Reading the bibliography…
A common explanation for the failure of out-of-distribution (OOD) generalization is that the model trained with empirical risk minimization (ERM) learns spurious features instead of invariant features.
The perceptron - a perceiving and recognizing automaton
F. Rosenblatt · 1957
Earlier work this paper cites.
Principles of risk minimization for learning theory
V. Vapnik · 1991
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Stochastic gradient methods for distributionally robust optimization with f-divergences
H. Namkoong and J. C. Duchi · 2016
Earlier work this paper cites.
Densely connected convolutional networks
G. Huang, Z. Liu, L. van der Maaten, and K. Q. Weinberger · 2017
Earlier work this paper cites.
Opening the black box of deep neural networks via information
R. Shwartz-Ziv and N. Tishby · 2017
Earlier work this paper cites.
Recognition in terra incognita
S. Beery, G. V. Horn, and P. Perona · 2018
Earlier work this paper cites.
SGD learns over-parameterized networks that provably generalize on linearly separable data
A. Brutzkus, A. Globerson, E. Malach, and S. Shalev-Shwartz · 2018
Earlier work this paper cites.
Functional map of the world
G. A. Christie, N. Fendley, J. Wilson, and R. Mukherjee · 2018
Earlier work this paper cites.
Invariant models for causal transfer learning
M. Rojas-Carulla, B. Schölkopf, R. Turner, and J. Peters · 2018
Earlier work this paper cites.
M. Arjovsky, L. Bottou, I. Gulrajani, and D. Lopez-Paz · 2019
Earlier work this paper cites.
From detection of individual metastases to classification of lymph node status at the patient level: The CAMELYON17 challenge
P. Bándi, O. Geessink, Q. Manson, M. V. Dijk, M. Balkenhol, M. Hermsen, B. E. Bejnordi, B. Lee, K. Paeng, A. Zhong, Q. Li, F. G. Zanjani, S. Zinger, K. Fukuta, D. Komura, V. Ovtcharov, S. Cheng, S. Zeng, J. Thagaard, A. B. Dahl, H. Lin, H. Chen, L. Jacobsson, M. Hedlund, M. Çetin, E. Halici, H. Jackson, R. Chen, F. Both, J. Franke, H. Küsters-Vandevelde, W. Vreuls, P. Bult, B. van Ginneken, J. van der Laak, and G. Litjens · 2019
Earlier work this paper cites.
Nuanced metrics for measuring unintended bias with real data for text classification
D. Borkan, L. Dixon, J. Sorensen, N. Thain, and L. Vasserman · 2019
Earlier work this paper cites.
SGD on neural networks learns functions of increasing complexity
D. Kalimeris, G. Kaplun, P. Nakkiran, B. Edelman, T. Yang, B. Barak, and H. Zhang · 2019
Earlier work this paper cites.
Justifying recommendations using distantly-labeled reviews and fine-grained aspects
J. Ni, J. Li, and J. McAuley · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala · 2019
Earlier work this paper cites.
Do ImageNet classifiers generalize to ImageNet?
B. Recht, R. Roelofs, L. Schmidt, and V. Shankar · 2019
Earlier work this paper cites.
Explainable AI: Interpreting, Explaining and Visualizing Deep Learning
W. Samek, G. Montavon, A. Vedaldi, L. K. Hansen, and K.-R. Muller · 2019
Earlier work this paper cites.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
V. Sanh, L. Debut, J. Chaumond, and T. Wolf · 2019
Earlier work this paper cites.
Rxrx1: An image set for cellular morphological variation across many experimental batches
J. Taylor, B. Earnshaw, B. Mabey, M. Victors, and J. Yosinski · 2019
Earlier work this paper cites.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Z. Allen-Zhu and Y. Li · 2020
Earlier work this paper cites.
The iwildcam 2020 competition dataset
S. Beery, E. Cole, and A. Gjoka · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Earlier work this paper cites.
Shortcut learning in deep neural networks
R. Geirhos, J. Jacobsen, C. Michaelis, R. S. Zemel, W. Brendel, M. Bethge, and F. A. Wichmann · 2020
Cited alongside, same era.
What shapes feature representations? exploring datasets, architectures, and training
K. Hermann and A. Lampinen · 2020
Cited alongside, same era.
The surprising simplicity of the early-time learning dynamics of neural networks
W. Hu, L. Xiao, B. Adlam, and J. Pennington · 2020
Cited alongside, same era.
Out-of-distribution generalization with maximal invariant predictor
M. Koyama and S. Yamaguchi · 2020
Cited alongside, same era.
Distributionally robust neural networks
S. Sagawa*, P. W. Koh*, T. B. Hashimoto, and P. Liang · 2020
Cited alongside, same era.
The pitfalls of simplicity bias in neural networks
Ensemble of averages: Improving model selection and boosting performance in domain generalization
D. Arpit, H. Wang, Y. Zhou, and C. Xiong · 2022
Later among the works it cites.
Benign overfitting in two-layer convolutional neural networks
Y. Cao, Z. Chen, M. Belkin, and Q. Gu · 2022
Later among the works it cites.
N. Elhage, T. Hume, C. Olsson, N. Schiefer, T. Henighan, S. Kravec, Z. Hatfield-Dodds, R. Lasenby, D. Drain, C. Chen, R. Grosse, S. McCandlish, J. Kaplan, D. Amodei, M. Wattenberg, and C. Olah · 2022
Later among the works it cites.
New definitions and evaluations for saliency methods: Staying intrinsic and sound, 2022
A. Gupta, N. Saunshi, D. Yu, K. Lyu, and S. Arora · 2022
Later among the works it cites.
On feature learning in the presence of spurious correlations
P. Izmailov, P. Kirichenko, N. Gruver, and A. G. Wilson · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Shah, K. Tamuly, A. Raghunathan, P. Jain, and P. Netrapalli · 2020
Cited alongside, same era.
Measuring robustness to natural distribution shifts in image classification
R. Taori, A. Dave, V. Shankar, N. Carlini, B. Recht, and L. Schmidt · 2020
Cited alongside, same era.
Invariance principle meets information bottleneck for out-of-distribution generalization
K. Ahuja, E. Caballero, D. Zhang, J.-C. Gagnon-Audet, Y. Bengio, I. Mitliagkas, and I. Rish · 2021
Cited alongside, same era.
The evolution of out-of-distribution robustness throughout fine-tuning
A. Andreassen, Y. Bahri, B. Neyshabur, and R. Roelofs · 2021
Cited alongside, same era.
AI for radiographic COVID-19 detection selects shortcuts over signal
A. J. DeGrave, J. D. Janizek, and S. Lee · 2021
Cited alongside, same era.
Provable generalization of sgd-trained neural networks of any width in the presence of adversarial label noise
S. Frei, Y. Cao, and Q. Gu · 2021
Cited alongside, same era.
In search of lost domain generalization
I. Gulrajani and D. Lopez-Paz · 2021
Cited alongside, same era.
Y. Ji, L. Zhang, J. Wu, B. Wu, L.-K. Huang, T. Xu, Y. Rong, L. Li, J. Ren, D. Xue, H. Lai, S. Xu, J. Feng, W. Liu, P. Luo, S. Zhou, J. Huang, P. Zhao, and Y. Bian · 2022
Later among the works it cites.
Last layer re-training is sufficient for robustness to spurious correlations
P. Kirichenko, P. Izmailov, and A. G. Wilson · 2022
Later among the works it cites.
Empirical study on optimizer selection for out-of-distribution generalization
H. Naganuma, K. Ahuja, I. Mitliagkas, S. Takagi, T. Motokawa, R. Yokota, K. Ishikawa, and I. Sato · 2022
Later among the works it cites.
Model ratatouille: Recycling diverse models for out-of-distribution generalization
A. Ramé, K. Ahuja, J. Zhang, M. Cord, L. Bottou, and D. Lopez-Paz · 2022
Later among the works it cites.
Diverse weight averaging for out-of-distribution generalization
A. Rame, M. Kirchmeyer, T. Rahier, A. Rakotomamonjy, patrick gallinari, and M. Cord · 2022
Later among the works it cites.
E. Rosenfeld, P. Ravikumar, and A. Risteski · 2022
Later among the works it cites.
Data augmentation as feature manipulation
R. Shen, S. Bubeck, and S. Gunasekar · 2022
Later among the works it cites.
Gradient matching for domain generalization
Y. Shi, J. Seely, P. Torr, S. N, A. Hannun, N. Usunier, and G. Synnaeve · 2022
Later among the works it cites.
Assaying out-of-distribution generalization in transfer learning
F. Wenzel, A. Dittadi, P. V. Gehler, C.-J. Simon-Gabriel, M. Horn, D. Zietlow, D. Kernert, C. Russell, T. Brox, B. Schiele, B. Schölkopf, and F. Locatello · 2022
Later among the works it cites.
Robust fine-tuning of zero-shot models
M. Wortsman, G. Ilharco, J. W. Kim, M. Li, S. Kornblith, R. Roelofs, R. G. Lopes, H. Hajishirzi, A. Farhadi, H. Namkoong, and L. Schmidt · 2022
Later among the works it cites.
H. Ye, J. Zou, and L. Zhang · 2022
Later among the works it cites.
Learning useful representations for shifting tasks and distributions
J. Zhang and L. Bottou · 2022
Later among the works it cites.
Model agnostic sample reweighting for out-of-distribution learning
X. Zhou, Y. Lin, R. Pi, W. Zhang, R. Xu, P. Cui, and T. Zhang · 2022
Later among the works it cites.
Robust learning with progressive data expansion against spurious correlation
Y. Deng, Y. Yang, B. Mirzasoleiman, and Q. Gu · 2023
Closest in time.
Graph neural networks provably benefit from structural information: A feature learning perspective
W. Huang, Y. Cao, H. Wang, X. Cao, and T. Suzuki · 2023
Closest in time.
Diversify and disambiguate: Out-of-distribution robustness via disagreement
Y. Lee, H. Yao, and C. Finn · 2023
Closest in time.
Towards stable backdoor purification through feature shift tuning
R. Min, Z. Qin, L. Shen, and M. Cheng · 2023
Closest in time.
Learning diverse features in vision transformers for improved generalization
A. M. Nicolicioiu, A. L. Nicolicioiu, B. Alexe, and D. Teney · 2023
Closest in time.
Towards out-of-distribution generalizable predictions of chemical kinetics properties
Z. Wang, Y. Chen, Y. Duan, W. Li, B. Han, J. Cheng, and H. Tong · 2023
Closest in time.