Fetching the paper…
Reading the bibliography…
Neural networks are known to exploit spurious artifacts (or shortcuts) that co-occur with a target label, exhibiting heuristic memorization.
Unmasking clever hans predictors and assessing what machines really learn
Lapuschkin, S., Wäldchen, S., Binder, A., Montavon, G., Samek, W., and Müller, K · 1902
Earlier work this paper cites.
Arjovsky, M., Bottou, L., Gulrajani, I., and Lopez-Paz, D · 1907
Earlier work this paper cites.
Roberta: A robustly optimized BERT pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 1907
Earlier work this paper cites.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 1910
Earlier work this paper cites.
Fantastic generalization measures and where to find them, 2019
Jiang, Y., Neyshabur, B., Mobahi, H., Krishnan, D., and Bengio, S · 1912
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Lecun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Estimation of the information by an adaptive partitioning of the observation space
Darbellay, G. and Vajda, I · 1999
Earlier work this paper cites.
The information bottleneck method
Tishby, N., Pereira, F. C., and Bialek, W · 1999
Earlier work this paper cites.
Estimating mutual information
Kraskov, A., Stögbauer, H., and Grassberger, P · 2004
Earlier work this paper cites.
A theory of learning from different domains
Ben-David, S., Blitzer, J., Crammer, K., Kulesza, A., Pereira, F., and Vaughan, J. W · 2010
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
An intuitive proof of the data processing inequality
Beaudry, N. J. and Renner, R · 2012
Earlier work this paper cites.
On causal and anticausal learning
Schölkopf, B., Janzing, D., Peters, J., Sgouritsa, E., Zhang, K., and Mooij, J. M · 2012
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2015
Earlier work this paper cites.
Deep learning and the information bottleneck principle
Tishby, N. and Zaslavsky, N · 2015
Earlier work this paper cites.
Probing for semantic evidence of composition by means of simple classification tasks
Ettinger, A., Elgohary, A., and Resnik, P · 2016
Earlier work this paper cites.
Fine-grained analysis of sentence embeddings using auxiliary prediction tasks
Adi, Y., Kermany, E., Belinkov, Y., Lavi, O., and Goldberg, Y · 2017
Earlier work this paper cites.
A closer look at memorization in deep networks
Arpit, D., Jastrzebski, S., Ballas, N., Krueger, D., Bengio, E., Kanwal, M. S., Maharaj, T., Fischer, A., Courville, A. C., Bengio, Y., and Lacoste-Julien, S · 2017
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations
Bau, D., Zhou, B., Khosla, A., Oliva, A., and Torralba, A · 2017
Earlier work this paper cites.
What do neural machine translation models learn about morphology?
Belinkov, Y., Durrani, N., Dalvi, F., Sajjad, H., and Glass, J. R · 2017
Cited alongside, same era.
Opening the black box of deep neural networks via information
Shwartz-Ziv, R. and Tishby, N · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2017
Cited alongside, same era.
Mutual information neural estimation
Belghazi, M. I., Baratin, A., Rajeswar, S., Ozair, S., Bengio, Y., Hjelm, R. D., and Courville, A. C · 2018
Cited alongside, same era.
Visualisation and ’diagnostic classifiers’ reveal how recurrent and recursive neural networks process hierarchical structure
Hupkes, D., Veldhoen, S., and Zuidema, W. H · 2018
Cited alongside, same era.
COGS: A compositional generalization challenge based on semantic interpretation
Kim, N. and Linzen, T · 2020
Later among the works it cites.
Neural complexity measures
Lee, Y., Lee, J., Hwang, S. J., Yang, E., and Choi, S · 2020
Later among the works it cites.
A constructive prediction of the generalization error across scales
Rosenfeld, J. S., Rosenfeld, A., Belinkov, Y., and Shavit, N · 2020
Later among the works it cites.
Dataset cartography: Mapping and diagnosing datasets with training dynamics
Swayamdipta, S., Schwartz, R., Lourie, N., Wang, Y., Hajishirzi, H., Smith, N. A., and Choi, Y · 2020
Later among the works it cites.
An Empirical Study on Robustness to Spurious Correlations using Pre-trained Language Models
Tu, L., Lalwani, G., Gella, S., and He, H · 2020
Later among the works it cites.
Information-theoretic probing with minimum description length
Voita, E. and Titov, I · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stress test evaluation for natural language inference
Naik, A., Ravichander, A., Sadeh, N. M., Rosé, C. P., and Neubig, G · 2018
Cited alongside, same era.
On the information bottleneck theory of deep learning
Saxe, A. M., Bansal, Y., Dapello, J., Advani, M., Kolchinsky, A., Tracey, B. D., and Cox, D. D · 2018
Cited alongside, same era.
Generating natural adversarial examples
Zhao, Z., Dua, D., and Singh, S · 2018
Cited alongside, same era.
What is one grain of sand in the desert? analyzing individual neurons in deep NLP models
Dalvi, F., Durrani, N., Sajjad, H., Belinkov, Y., Bau, A., and Glass, J. R · 2019
Cited alongside, same era.
Bias in bios: A case study of semantic representation bias in a high-stakes setting
De-Arteaga, M., Romanov, A., Wallach, H. M., Chayes, J. T., Borgs, C., Chouldechova, A., Geyik, S. C., Kenthapadi, K., and Kalai, A. T · 2019
Cited alongside, same era.
Imagenet-trained cnns are biased towards texture; increasing shape bias improves accuracy and robustness
Geirhos, R., Rubisch, P., Michaelis, C., Bethge, M., Wichmann, F. A., and Brendel, W · 2019
Cited alongside, same era.
Benchmarking neural network robustness to common corruptions and perturbations
Hendrycks, D. and Dietterich, T. G · 2019
Cited alongside, same era.
Later among the works it cites.
Types of out-of-distribution texts and how to detect them
Arora, U., Huang, W., and He, H · 2021
Later among the works it cites.
The many faces of robustness: A critical analysis of out-of-distribution generalization
Hendrycks, D., Basart, S., Mu, N., Kadavath, S., Wang, F., Dorundo, E., Desai, R., Zhu, T., Parajuli, S., Guo, M., Song, D., Steinhardt, J., and Gilmer, J · 2021
Later among the works it cites.
Natural adversarial examples
Hendrycks, D., Zhao, K., Basart, S., Steinhardt, J., and Song, D · 2021
Later among the works it cites.
Variational information bottleneck for effective low-resource fine-tuning
Mahabadi, R. K., Belinkov, Y., and Henderson, J · 2021
Later among the works it cites.
NoiseQA: Challenge set evaluation for user-centric question answering
Ravichander, A., Dalmia, S., Ryskina, M., Metze, F., Hovy, E., and Black, A. W · 2021
Later among the works it cites.
Towards out-of-distribution generalization: A survey
Shen, Z., Liu, J., He, Y., Zhang, X., Xu, R., Yu, H., and Cui, P · 2021
Later among the works it cites.
BERT memorisation and pitfalls in low-resource scenarios
Tänzer, M., Ruder, S., and Rei, M · 2021
Later among the works it cites.
Infobert: Improving robustness of language models from an information theoretic perspective
Wang, B., Wang, S., Cheng, Y., Gan, Z., Jia, R., Li, B., and Liu, J · 2021
Later among the works it cites.
Generalizing to unseen domains: A survey on domain generalization
Wang, J., Lan, C., Liu, C., Ouyang, Y., and Qin, T · 2021
Later among the works it cites.
On the pitfalls of analyzing individual neurons in language models
Antverg, O. and Belinkov, Y · 2022
Closest in time.
Probing classifiers: Promises, shortcomings, and advances
Belinkov, Y · 2022
Closest in time.
How gender debiasing affects internal model representations, and why it matters
Orgad, H., Goldfarb-Tarrant, S., and Belinkov, Y · 2022
Closest in time.
Nico++: Towards better benchmarking for domain generalization, 2022
Zhang, X., He, Y., Xu, R., Yu, H., Shen, Z., and Cui, P · 2022
Closest in time.