Fetching the paper…
Reading the bibliography…
Deep classifiers are known to rely on spurious features $\unicode{x2013}$ patterns which are correlated with the target on the training data but not inherently relevant to the learning problem, such as the image backgrounds when classifying the foregrounds.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
McCoy, R. T., Pavlick, E., and Linzen, T. (2019) · 1902
Earlier work this paper cites.
Approximating cnns with bag-of-local-features models works surprisingly well on imagenet
Brendel, W. and Bethge, M. (2019) · 1904
Earlier work this paper cites.
Maximum weighted loss discrepancy
Khani, F., Raghunathan, A., and Liang, P. (2019) · 1906
Earlier work this paper cites.
Arjovsky, M., Bottou, L., Gulrajani, I., and Lopez-Paz, D. (2019) · 1907
Earlier work this paper cites.
Probing neural network comprehension of natural language arguments
Niven, T. and Kao, H.-Y. (2019) · 1907
Earlier work this paper cites.
Decoupling representation and classifier for long-tailed recognition
Kang, B., Xie, S., Rohrbach, M., Yan, Z., Gordo, A., Feng, J., and Kalantidis, Y. (2019) · 1910
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al. (2019) · 1910
Earlier work this paper cites.
Sagawa, S., Koh, P. W., Hashimoto, T. B., and Liang, P. (2019) · 1911
Earlier work this paper cites.
Augmix: A simple data processing method to improve robustness and uncertainty
Hendrycks, D., Mu, N., Cubuk, E. D., Zoph, B., Gilmer, J., and Lakshminarayanan, B. (2019) · 1912
Earlier work this paper cites.
A stochastic approximation method
Robbins, H. and Monro, S. (1951) · 1951
Earlier work this paper cites.
Deberta: Decoding-enhanced bert with disentangled attention
He, P., Liu, X., Gao, J., and Chen, W. (2020) · 2006
Earlier work this paper cites.
Noise or signal: The role of image backgrounds in object recognition
Xiao, K., Engstrom, L., Ilyas, A., and Madry, A. (2020) · 2006
Earlier work this paper cites.
Matplotlib: A 2d graphics environment
Hunter, J. D. (2007) · 2007
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. (2020) · 2010
Earlier work this paper cites.
Torchvision the machine-vision package of torch
Marcel, S. and Rodriguez, Y. (2010) · 2010
Earlier work this paper cites.
Data structures for statistical computing in python
McKinney, W. (2010) · 2010
Earlier work this paper cites.
The caltech-ucsd birds-200-2011 dataset
Wah, C., Branson, S., Welinder, P., Perona, P., and Belongie, S. (2011) · 2011
Earlier work this paper cites.
Fairness through awareness
Dwork, C., Hardt, M., Pitassi, T., Reingold, O., and Zemel, R. (2012) · 2012
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Krizhevsky, A. (2014) · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A. (2014) · 2014
Earlier work this paper cites.
Deep learning face attributes in the wild
Liu, Z., Luo, P., Wang, X., and Tang, X. (2015) · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al. (2015) · 2015
Earlier work this paper cites.
Equality of opportunity in supervised learning
Hardt, M., Price, E., and Srebro, N. (2016) · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Earlier work this paper cites.
Inherent trade-offs in the fair determination of risk scores
Kleinberg, J., Mullainathan, S., and Raghavan, M. (2016) · 2016
Earlier work this paper cites.
Jupyter notebooks – a publishing format for reproducible computational workflows
Kluyver, T., Ragan-Kelley, B., Pérez, F., Granger, B., Bussonnier, M., Frederic, J., Kelley, K., Hamrick, J., Grout, J., Corlay, S., Ivanov, P., Avila, D., Abdalla, S., and Willing, C. (2016) · 2016
Earlier work this paper cites.
Improving weakly-supervised object localization by micro-annotation
Kolesnikov, A. and Lampert, C. H. (2016) · 2016
Earlier work this paper cites.
Zagoruyko, S. and Komodakis, N. (2016) · 2016
Earlier work this paper cites.
Densely connected convolutional networks
Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q. (2017) · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B., Pritzel, A., and Blundell, C. (2017) · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F. (2017) · 2017
Earlier work this paper cites.
Automatic differentiation in pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A. (2017) · 2017
Earlier work this paper cites.
On fairness and calibration
Pleiss, G., Raghavan, M., Wu, F., Kleinberg, J., and Weinberger, K. Q. (2017) · 2017
Earlier work this paper cites.
Chexnet: Radiologist-level pneumonia detection on chest x-rays with deep learning
Rajpurkar, P., Irvin, J., Zhu, K., Yang, B., Mehta, H., Duan, T., Ding, D., Bagul, A., Langlotz, C., Shpanskaya, K., et al. (2017) · 2017
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A., Nangia, N., and Bowman, S. R. (2017) · 2017
Cited alongside, same era.
Aggregated residual transformations for deep neural networks
Xie, S., Girshick, R., Dollár, P., Tu, Z., and He, K. (2017) · 2017
Cited alongside, same era.
mixup: Beyond empirical risk minimization
Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D. (2017) · 2017
Cited alongside, same era.
Places: A 10 million image database for scene recognition
Zhou, B., Lapedriza, A., Khosla, A., Oliva, A., and Torralba, A. (2017) · 2017
SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python
Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van der Walt, S. J., Brett, M., Wilson, J., Millman, K. J., Mayorov, N., Nelson, A. R. J., Jones, E., Kern, R., Larson, E., Carey, C. J., Polat, İ., Feng, Y., Moore, E. W., VanderPlas, J., Laxalde, D., Perktold, J., Cimrman, R., Henriksen, I., Quintero, E. A., Harris, C. R., Archibald, A. M., Ribeiro, A. H., Pedregosa, F., van Mulbregt, P., and SciPy 1.0 Contributors (2020) · 2020
Later among the works it cites.
Random erasing data augmentation
Zhong, Z., Zheng, L., Kang, G., Li, S., and Yang, Y. (2020) · 2020
Later among the works it cites.
Beit: Bert pre-training of image transformers
Bao, H., Dong, L., and Wei, F. (2021) · 2021
Later among the works it cites.
Emerging properties in self-supervised vision transformers
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., and Joulin, A. (2021) · 2021
Later among the works it cites.
Environment inference for invariant learning
Creager, E., Jacobsen, J.-H., and Zemel, R. (2021) · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A reductions approach to fair classification
Agarwal, A., Beygelzimer, A., Dudík, M., Langford, J., and Wallach, H. (2018) · 2018
Cited alongside, same era.
Functional map of the world
Christie, G., Fendley, N., Wilson, J., and Mukherjee, R. (2018) · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018) · 2018
Cited alongside, same era.
Geirhos, R., Rubisch, P., Michaelis, C., Bethge, M., Wichmann, F. A., and Brendel, W. (2018) · 2018
Cited alongside, same era.
Annotation artifacts in natural language inference data
Gururangan, S., Swayamdipta, S., Levy, O., Schwartz, R., Bowman, S. R., and Smith, N. A. (2018) · 2018
Cited alongside, same era.
How much reading does reading comprehension require? a critical investigation of popular benchmarks
Kaushik, D. and Lipton, Z. C. (2018) · 2018
Cited alongside, same era.
Resound: Towards action recognition without representation bias
Li, Y., Li, Y., and Vasconcelos, N. (2018) · 2018
Cited alongside, same era.
Later among the works it cites.
Simple data balancing achieves competitive worst-group-accuracy
Idrissi, B. Y., Arjovsky, M., Pezeshki, M., and Lopez-Paz, D. (2021) · 2021
Later among the works it cites.
Explaining the efficacy of counterfactually augmented data
Kaushik, D., Setlur, A., Hovy, E., and Lipton, Z. C. (2021) · 2021
Later among the works it cites.
Wilds: A benchmark of in-the-wild distribution shifts
Koh, P. W., Sagawa, S., Marklund, H., Xie, S. M., Zhang, M., Balsubramani, A., Hu, W., Yasunaga, M., Phillips, R. L., Gao, I., et al. (2021) · 2021
Later among the works it cites.
Just train twice: Improving group robustness without training group information
Liu, E. Z., Haghgoo, B., Chen, A. S., Raghunathan, A., Koh, P. W., Sagawa, S., Liang, P., and Finn, C. (2021) · 2021
Later among the works it cites.
Accuracy on the line: on the strong correlation between out-of-distribution and in-distribution generalization
Miller, J. P., Taori, R., Raghunathan, A., Sagawa, S., Koh, P. W., Shankar, V., Liang, P., Carmon, Y., and Schmidt, L. (2021) · 2021
Later among the works it cites.
Object-aware contrastive learning for debiased scene representation
Mo, S., Kang, H., Sohn, K., Li, C.-L., and Shin, J. (2021) · 2021
Later among the works it cites.
Gradient starvation: A learning proclivity in neural networks
Pezeshki, M., Kaba, O., Bengio, Y., Courville, A. C., Precup, D., and Lajoie, G. (2021) · 2021
Later among the works it cites.
Optimal representations for covariate shift
Ruan, Y., Dubois, Y., and Maddison, C. J. (2021) · 2021
Later among the works it cites.
Extending the wilds benchmark for unsupervised adaptation
Sagawa, S., Koh, P. W., Lee, T., Gao, I., Xie, S. M., Shen, K., Kumar, A., Hu, W., Yasunaga, M., Marklund, H., et al. (2021) · 2021
Later among the works it cites.
Salient imagenet: How to discover spurious features in deep learning?
Singla, S. and Feizi, S. (2021) · 2021
Later among the works it cites.
Barack: Partially supervised group robustness with guarantees
Sohoni, N., Sanjabi, M., Ballas, N., Grover, A., Nie, S., Firooz, H., and Ré, C. (2021) · 2021
Later among the works it cites.
Teney, D., Abbasnejad, E., Lucey, S., and Hengel, A. v. d. (2021) · 2021
Later among the works it cites.
Counterfactual invariance to spurious correlations: Why and how to pass stress tests
Veitch, V., D’Amour, A., Yadlowsky, S., and Eisenstein, J. (2021) · 2021
Later among the works it cites.
Barlow twins: Self-supervised learning via redundancy reduction
Zbontar, J., Jing, L., Misra, I., LeCun, Y., and Deny, S. (2021) · 2021
Later among the works it cites.
Correct-n-contrast: A contrastive approach for improving robustness to spurious correlations
Zhang, M., Sohoni, N. S., Zhang, H. R., Finn, C., and Ré, C. (2021) · 2021
Later among the works it cites.
The effects of regularization and data augmentation are class dependent
Balestriero, R., Bottou, L., and LeCun, Y. (2022) · 2022
Closest in time.
Informativeness and invariance: Two perspectives on spurious correlations in natural language
Eisenstein, J. (2022) · 2022
Closest in time.
Are vision transformers robust to spurious correlations?
Ghosal, S. S., Ming, Y., and Li, Y. (2022) · 2022
Closest in time.
Last layer re-training is sufficient for robustness to spurious correlations
Kirichenko, P., Izmailov, P., and Wilson, A. G. (2022) · 2022
Closest in time.
Diversify and disambiguate: Learning from underspecified data
Lee, Y., Yao, H., and Finn, C. (2022) · 2022
Closest in time.
Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., and Xie, S. (2022) · 2022
Closest in time.
Moayeri, M., Pope, P., Balaji, Y., and Feizi, S. (2022) · 2022
Closest in time.
Spread spurious attribute: Improving worst-group accuracy with spurious attribute estimation
Nam, J., Kim, J., Lee, J., and Shin, J. (2022) · 2022
Closest in time.
Agree to disagree: Diversity through disagreement for better transferability
Pagliardini, M., Jaggi, M., Fleuret, F., and Karimireddy, S. P. (2022) · 2022
Closest in time.
Rosenfeld, E., Ravikumar, P., and Risteski, A. (2022) · 2022
Closest in time.
How robust are pre-trained models to distribution shift?
Shi, Y., Daunhawer, I., Vogt, J. E., Torr, P. H., and Sanyal, A. (2022) · 2022
Closest in time.
Rich feature construction for the optimization-generalization dilemma
Zhang, J., Lopez-Paz, D., and Bottou, L. (2022) · 2022
Closest in time.
Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases
Wang, X., Peng, Y., Lu, L., Lu, Z., Bagheri, M., and Summers, R. M. (2017) · 2097
Closest in time.