Fetching the paper…
Reading the bibliography…
Out-of-distribution generalization in neural networks is often hampered by spurious correlations.
Arjovsky, M., Bottou, L., Gulrajani, I., and Lopez-Paz, D · 1907
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., and Brew, J · 1910
Earlier work this paper cites.
On the asymptotic distribution of the sum of a random number of independent random variables
Rényi, A · 1957
Earlier work this paper cites.
Caltech-ucsd birds 200
Welinder, P., Branson, S., Mita, T., Wah, C., Schroff, F., Belongie, S., and Perona, P · 2010
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Unsupervised domain adaptation by backpropagation
Ganin, Y. and Lempitsky, V · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Deep learning face attributes in the wild
Liu, Z., Luo, P., Wang, X., and Tang, X · 2015
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Bolukbasi, T., Chang, K.-W., Zou, J. Y., Saligrama, V., and Kalai, A. T · 2016
Earlier work this paper cites.
Censoring representations with an adversary
Edwards, H. and Storkey, A · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Places: An image database for deep scene understanding
Zhou, B., Khosla, A., Lapedriza, À., Torralba, A., and Oliva, A · 2016
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations
Bau, D., Zhou, B., Khosla, A., Oliva, A., and Torralba, A · 2017
Earlier work this paper cites.
Annotation artifacts in natural language inference data
Gururangan, S., Swayamdipta, S., Levy, O., Schwartz, R., Bowman, S., and Smith, N. A · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D · 2017
Earlier work this paper cites.
Cleaning the null space: a privacy mechanism for predictors
Xu, K., Cao, T., Shah, S., Maung, C., and Schweitzer, H · 2017
Earlier work this paper cites.
Adversarial removal of demographic attributes from text data
Elazar, Y. and Goldberg, Y · 2018
Earlier work this paper cites.
How much reading does reading comprehension require? a critical investigation of popular benchmarks
Kaushik, D. and Lipton, Z. C · 2018
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV)
Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., Viegas, F., and sayres, R · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A., Nangia, N., and Bowman, S · 2018
Earlier work this paper cites.
Mitigating unwanted biases with adversarial learning
Zhang, B. H., Lemoine, B., and Mitchell, M · 2018
Earlier work this paper cites.
Attenuating bias in word vectors
Dev, S. and Phillips, J · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Understanding undesirable word embedding associations
Ethayarajh, K., Duvenaud, D., and Hirst, G · 2019
Cited alongside, same era.
Imagenet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness
Geirhos, R., Rubisch, P., Michaelis, C., Bethge, M., Wichmann, F. A., and Brendel, W · 2019
Cited alongside, same era.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
McCoy, T., Pavlick, E., and Linzen, T · 2019
Cited alongside, same era.
Probing neural network comprehension of natural language arguments
Niven, T. and Kao, H.-Y · 2019
Probing the probing paradigm: Does probing accuracy entail task relevance?
Ravichander, A., Belinkov, Y., and Hovy, E · 2021
Later among the works it cites.
Dynamically disentangling social bias from task-oriented representations with adversarial attack
Wang, L., Yan, Y., He, K., Wu, Y., and Xu, W · 2021
Later among the works it cites.
Noise or signal: The role of image backgrounds in object recognition
Xiao, K. Y., Engstrom, L., Ilyas, A., and Madry, A · 2021
Later among the works it cites.
Disentangling representations of text by masking transformers
Zhang, X., van de Meent, J.-W., and Wallace, B · 2021
Later among the works it cites.
Examining and combating spurious features under distribution shift
Zhou, C., Ma, X., Michel, P., and Neubig, G · 2021
Later among the works it cites.
Probing classifiers: Promises, shortcomings, and advances
Belinkov, Y · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the spectral bias of neural networks
Rahaman, N., Baratin, A., Arpit, D., Draxler, F., Lin, M., Hamprecht, F., Bengio, Y., and Courville, A · 2019
Cited alongside, same era.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Rudin, C · 2019
Cited alongside, same era.
Balanced datasets are not enough: Estimating and mitigating gender bias in deep image representations
Wang, T., Zhao, J., Yatskar, M., Chang, K.-W., and Ordonez, V · 2019
Cited alongside, same era.
Concept whitening for interpretable image recognition
Chen, Z., Bei, Y., and Rudin, C · 2020
Cited alongside, same era.
The origins and prevalence of texture bias in convolutional neural networks
Hermann, K., Chen, T., and Kornblith, S · 2020
Cited alongside, same era.
Null it out: Guarding protected attributes by iterative nullspace projection
Ravfogel, S., Elazar, Y., Gonen, H., Twiton, M., and Goldberg, Y · 2020
Cited alongside, same era.
Later among the works it cites.
Discovering latent concepts learned in BERT
Dalvi, F., Khan, A. R., Alam, F., Durrani, N., Xu, J., and Sajjad, H · 2022
Later among the works it cites.
Underspecification presents challenges for credibility in modern machine learning
D’Amour, A., Heller, K., Moldovan, D., Adlam, B., Alipanahi, B., Beutel, A., Chen, C., Deaton, J., Eisenstein, J., Hoffman, M. D., Hormozdiari, F., Houlsby, N., Hou, S., Jerfel, G., Karthikesalingam, A., Lucic, M., Ma, Y., McLean, C., Mincu, D., Mitani, A., Montanari, A., Nado, Z., Natarajan, V., Nielson, C., Osborne, T. F., Raman, R., Ramasamy, K., Sayres, R., Schrouff, J., Seneviratne, M., Sequeira, S., Suresh, H., Veitch, V., Vladymyrov, M., Wang, X., Webster, K., Yadlowsky, S., Yun, T., Zhai, X., and Sculley, D · 2022
Later among the works it cites.
Simple data balancing achieves competitive worst-group-accuracy
Idrissi, B. Y., Arjovsky, M., Pezeshki, M., and Lopez-Paz, D · 2022
Later among the works it cites.
On feature learning in the presence of spurious correlations
Izmailov, P., Kirichenko, P., Gruver, N., and Wilson, A. G · 2022
Later among the works it cites.
Are all spurious features in natural language alike? an analysis through a causal lens
Joshi, N., Pan, X., and He, H · 2022
Later among the works it cites.
Probing classifiers are unreliable for concept removal and detection
Kumar, A., Tan, C., and Sharma, A · 2022
Later among the works it cites.
Adversarial concept erasure in kernel space
Ravfogel, S., Vargas, F., Goldberg, Y., and Cotterell, R · 2022
Later among the works it cites.
Domain-adjusted regression or: ERM may already learn features sufficient for out-of-distribution generalization
Rosenfeld, E., Ravikumar, P. K., and Risteski, A · 2022
Later among the works it cites.
Interpretable machine learning: Fundamental principles and 10 grand challenges
Rudin, C., Chen, C., Chen, Z., Huang, H., Semenova, L., and Zhong, C · 2022
Later among the works it cites.
Salient imagenet: How to discover spurious features in deep learning?
Singla, S. and Feizi, S · 2022
Later among the works it cites.
Leace: Perfect linear concept erasure in closed form
Belrose, N., Schneider-Joseph, D., Ravfogel, S., Cotterell, R., Raff, E., and Biderman, S · 2023
Closest in time.
Last layer re-training is sufficient for robustness to spurious correlations
Kirichenko, P., Izmailov, P., and Wilson, A. G · 2023
Closest in time.
The linear representation hypothesis and the geometry of large language models
Park, K., Choe, Y. J., and Veitch, V · 2023
Closest in time.
Log-linear guardedness and its implications
Ravfogel, S., Goldberg, Y., and Cotterell, R · 2023
Closest in time.