Fetching the paper…
Reading the bibliography…
While deep learning models have shown remarkable performance in various tasks, they are susceptible to learning non-generalizable spurious features rather than the core features that are genuinely correlated to the true label.
Improving predictive inference under covariate shift by weighting the log-likelihood function
Shimodaira, H · 2000
Earlier work this paper cites.
Visualizing data using t-sne
Van der Maaten, L. and Hinton, G · 2008
Earlier work this paper cites.
Learning from imbalanced data
He, H. and Garcia, E. A · 2009
Earlier work this paper cites.
The caltech-ucsd birds-200-2011 dataset
Wah, C., Branson, S., Welinder, P., Perona, P., and Belongie, S · 2011
Earlier work this paper cites.
Deep learning face attributes in the wild
Liu, Z., Luo, P., Wang, X., and Tang, X · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Variance-based regularization with convex objectives
Namkoong, H. and Duchi, J. C · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Places: A 10 million image database for scene recognition
Zhou, B., Lapedriza, A., Khosla, A., Oliva, A., and Torralba, A · 2017
Earlier work this paper cites.
Does distributionally robust supervised learning give robust classifiers?
Hu, W., Niu, G., Sato, I., and Sugiyama, M · 2018
Earlier work this paper cites.
What is the effect of importance weighting in deep learning?
Byrd, J. and Lipton, Z · 2019
Earlier work this paper cites.
Learning imbalanced datasets with label-distribution-aware margin loss
Cao, K., Wei, C., Gaidon, A., Arechiga, N., and Ma, T · 2019
Earlier work this paper cites.
Class-balanced loss based on effective number of samples
Cui, Y., Jia, M., Lin, T.-Y., Song, Y., and Belongie, S · 2019
Earlier work this paper cites.
Distributionally robust losses against mixture covariate shifts
Duchi, J. C., Hashimoto, T., and Namkoong, H · 2019
Earlier work this paper cites.
Distributionally robust language modeling
Oren, Y., Sagawa, S., Hashimoto, T. B., and Liang, P · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Cited alongside, same era.
Distributionally robust neural networks
Sagawa, S., Koh, P. W., Hashimoto, T. B., and Liang, P · 2019
Cited alongside, same era.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Allen-Zhu, Z. and Li, Y · 2020
Cited alongside, same era.
Heteroskedastic and imbalanced deep learning with adaptive regularization
Cao, K., Chen, Y., Lu, J., Arechiga, N., Gaidon, A., and Ma, T · 2020
Cited alongside, same era.
Self-training avoids using spurious features under domain shift
Chen, Y., Wei, C., Kumar, A., and Ma, T · 2020
Cited alongside, same era.
Coping with label shift via distributionally robust optimisation
Zhang, J., Menon, A. K., Veit, A., Bhojanapalli, S., Kumar, S., and Sra, S · 2021
Later among the works it cites.
Understanding the generalization of adam in learning neural networks with proper regularization
Zou, D., Cao, Y., Li, Y., and Gu, Q · 2021
Later among the works it cites.
Benign overfitting in two-layer convolutional neural networks
Cao, Y., Chen, Z., Belkin, M., and Gu, Q · 2022
Later among the works it cites.
Towards understanding the mixture-of-experts layer in deep learning
Chen, Z., Deng, Y., Wu, Y., Gu, Q., and Li, Y · 2022
Later among the works it cites.
On-demand sampling: Learning optimally from multiple distributions
Haghtalab, N., Jordan, M., and Zhao, E · 2022
Later among the works it cites.
On feature learning in the presence of spurious correlations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning from failure: De-biasing classifier from biased classifier
Nam, J., Cha, H., Ahn, S., Lee, J., and Shin, J · 2020
Cited alongside, same era.
An investigation of why overparameterization exacerbates spurious correlations
Sagawa, S., Raghunathan, A., Koh, P. W., and Liang, P · 2020
Cited alongside, same era.
No subclass left behind: Fine-grained robustness in coarse-grained classification problems
Sohoni, N., Dunnmon, J., Angus, G., Gu, A., and Ré, C · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M · 2020
Cited alongside, same era.
Systematic generalisation with group invariant predictions
Ahmed, F., Bengio, Y., Van Seijen, H., and Courville, A · 2021
Cited alongside, same era.
Environment inference for invariant learning
Creager, E., Jacobsen, J.-H., and Zemel, R · 2021
Cited alongside, same era.
Model patching: Closing the subgroup performance gap with data augmentation
Goel, K., Gu, A., Li, Y., and Re, C · 2021
Cited alongside, same era.
Izmailov, P., Kirichenko, P., Gruver, N., and Wilson, A. G · 2022
Later among the works it cites.
Towards understanding how momentum improves generalization in deep learning
Jelassi, S. and Li, Y · 2022
Later among the works it cites.
Understanding rare spurious correlations in neural networks
Yang, Y.-Y., Chou, C.-N., and Chaudhuri, K · 2022
Later among the works it cites.
Ye, H., Zou, J., and Zhang, L · 2022
Later among the works it cites.
Correct-n-contrast: A contrastive approach for improving robustness to spurious correlations
Zhang, M., Sohoni, N. S., Zhang, H. R., Finn, C., and Ré, C · 2022
Later among the works it cites.
Towards understanding feature learning in out-of-distribution generalization
Chen, Y., Huang, W., Zhou, K., Bian, Y., Han, B., and Cheng, J · 2023
Closest in time.
Towards mitigating spurious correlations in the wild: A benchmark & a more realistic dataset
Joshi, S., Yang, Y., Xue, Y., Yang, W., and Mirzasoleiman, B · 2023
Closest in time.
Last layer re-training is sufficient for robustness to spurious correlations
Kirichenko, P., Izmailov, P., and Wilson, A. G · 2023
Closest in time.
Distributionally robust post-hoc classifiers under prior shifts
Wei, J., Narasimhan, H., Amid, E., Chu, W.-S., Liu, Y., and Kumar, A · 2023
Closest in time.
Eliminating spurious correlations from pre-trained models via data mixing
Xue, Y., Payani, A., Yang, Y., and Mirzasoleiman, B · 2023
Closest in time.