Fetching the paper…
Reading the bibliography…
We describe an "interpretability illusion" that arises when analyzing the BERT model.
Manzini, T., Lim, Y. C., Tsvetkov, Y., and Black, A. W · 1904
Earlier work this paper cites.
Gender-preserving debiasing for pre-trained word embeddings
Kaneko, M. and Bollegala, D · 1906
Earlier work this paper cites.
A neural probabilistic language model
Bengio, Y., Ducharme, R., Vincent, P., and Janvin, C · 2003
Earlier work this paper cites.
Analyzing individual neurons in pre-trained language models
Durrani, N., Dalvi, F., Sajjad, H., and Belinkov, Y · 2010
Earlier work this paper cites.
Efficient estimation of word representations in vector space, 2013
Mikolov, T., Chen, K., Corrado, G., and Dean, J · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Zeiler, M. D. and Fergus, R · 2014
Earlier work this paper cites.
Visualizing and understanding neural models in NLP
Li, J., Chen, X., Hovy, E. H., and Jurafsky, D · 2015
Earlier work this paper cites.
Object detectors emerge in deep scene cnn
Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., and Torralba, A · 2015
Earlier work this paper cites.
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Bolukbasi, T. and Chang, K.-W · 2016
Cited alongside, same era.
Synthesizing the preferred inputs for neurons in neural networks via deep generator networks
Nguyen, A., Dosovitskiy, A., Yosinski, J., Brox, T., and Clune, J · 2016
Cited alongside, same era.
Network dissection: Quantifying interpretability of deep visual representations
Bau, D., Zhou, B., Khosla, A., Oliva, A., and Torralba, A · 2017
Cited alongside, same era.
First quora dataset release: Question pairs, 2017
Iyer, S., Dandekar, N., and Csernai, K · 2017
Cited alongside, same era.
Feature visualization
Olah, C., Mordvintsev, A., and Schubert, L · 2017
Cited alongside, same era.
Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability, 2017
Raghu, M., Gilmer, J., Yosinski, J., and Sohl-Dickstein, J · 2017
What is one grain of sand in the desert? analyzing individual neurons in deep nlp models
Dalvi, F., Durrani, N., and Sajjad, H · 2019
Later among the works it cites.
Discovery of natural language concepts in individual units of cnns
Na, S., Choe, Y. J., Lee, D.-H., and Kim, G · 2019
Later among the works it cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Reimers, N. and Gurevych, I · 2019
Later among the works it cites.
What do you learn from context? probing for sentence structure in contextualized word representations
Tenney, I., Xia, P., Chen, B., Wang, A., Poliak, A., McCoy, R. T., Kim, N., Durme, B. V., Bowman, S., Das, D., and Pavlick, E · 2019
Later among the works it cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., Viegas, F., et al · 2018
Cited alongside, same era.
Interpretable textual neuron representations for NLP
Poerner, N., Roth, B., and Schütze, H · 2018
Cited alongside, same era.
What does bert dream of?
Bäuerle, A. and Wexler, J
Cited in the paper.
Aharoni, R. and Goldberg, Y · 2020
Later among the works it cites.
Emergent linguistic structure in artificial neural networks trained by self-supervision
Manning, C., Hewitt, J., Clark, K., Khandelwal, U., and Levy, O · 2020
Later among the works it cites.
Zoom in: An introduction to circuits
Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S · 2020
Later among the works it cites.
Causal mediation analysis for interpreting neural nlp: The case of gender bias, 2020
Vig, J., Gehrmann, S., Belinkov, Y., Qian, S., Nevo, D., Sakenis, S., Huang, J., Singer, Y., and Shieber, S · 2020
Later among the works it cites.