Fetching the paper…
Reading the bibliography…
Common methods for interpreting neural models in natural language processing typically examine either their structure or their behavior, but not both.
HuggingFace’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., and Brew, J. (2019) · 1910
Earlier work this paper cites.
Best algorithms for approximating the maximum of a submodular set function
Nemhauser, G. L., and Wolsey, L. A. (1978) · 1978
Earlier work this paper cites.
Statistics and causal inference
Holland, P. W. (1986) · 1986
Earlier work this paper cites.
Identifiability and exchangeability for direct and indirect effects
Robins, J. M., and Greenland, S. (1992) · 1992
Earlier work this paper cites.
Causal diagrams for empirical research
Pearl, J. (1995) · 1995
Earlier work this paper cites.
Causation, prediction, and search
Spirtes, P., Glymour, C. N., Scheines, R., and Heckerman, D. (2000) · 2000
Earlier work this paper cites.
Direct and indirect effects
Pearl, J. (2001) · 2001
Earlier work this paper cites.
Semantics of causal DAG models and the identification of direct and indirect effects
Robins, J. M. (2003) · 2003
Earlier work this paper cites.
Identifiability of path-specific effects
Avin, C., Shpitser, I., and Pearl, J. (2005) · 2005
Earlier work this paper cites.
Bootstrapping path-based pronoun resolution
Bergsma, S., and Lin, D. (2006) · 2006
Earlier work this paper cites.
Causality
Pearl, J. (2009) · 2009
Earlier work this paper cites.
Conceptual issues concerning mediation, interventions and composition
VanderWeele, T. J., and Vansteelandt, S. (2009) · 2009
Earlier work this paper cites.
Fairness through awareness.
Dwork, C., Hardt, M., Pitassi, T., Reingold, O., and Zemel, R. (2011) · 2011
Earlier work this paper cites.
Submodular maximization with cardinality constraints
Buchbinder, N., Feldman, M., Naor, J. S., and Schwartz, R. (2014) · 2014
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Bolukbasi, T., Chang, K.-W., Zou, J. Y., Saligrama, V., and Kalai, A. T. (2016) · 2016
Earlier work this paper cites.
Visualizing and understanding neural models in NLP
Li, J., Chen, X., Hovy, E., and Jurafsky, D. (2016) · 2016
Earlier work this paper cites.
Fine-grained analysis of sentence embeddings using auxiliary prediction tasks
Adi, Y., Kermany, E., Belinkov, Y., Lavi, O., and Goldberg, Y. (2017) · 2017
Earlier work this paper cites.
What is relevant in a text document?”: An interpretable machine learning approach
Arras, L., Horn, F., Montavon, G., Müller, K.-R., and Samek, W. (2017) · 2017
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Caliskan, A., Bryson, J. J., and Narayanan, A. (2017) · 2017
Earlier work this paper cites.
A challenge set approach to evaluating machine translation
Isabelle, P., Cherry, C., and Foster, G. (2017) · 2017
Earlier work this paper cites.
Counterfactual fairness
Kusner, M. J., Loftus, J., Russell, C., and Silva, R. (2017) · 2017
Earlier work this paper cites.
Social bias in elicited natural language inferences
Rudinger, R., May, C., and Van Durme, B. (2017) · 2017
Earlier work this paper cites.
How grammatical is character-level neural machine translation? assessing MT quality with contrastive translation pairs
Sennrich, R. (2017) · 2017
Earlier work this paper cites.
Lstmvis: A tool for visual analysis of hidden state dynamics in recurrent neural networks
Strobelt, H., Gehrmann, S., Pfister, H., and Rush, A. M. (2017) · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
Non-monotone submodular maximization in exponentially fewer iterations
Balkanski, E., Breuer, A., and Singer, Y. (2018) · 2018
Earlier work this paper cites.
The adaptive complexity of maximizing a submodular function
Balkanski, E., and Singer, Y. (2018a) · 2018
Earlier work this paper cites.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Conneau, A., Kruszewski, G., Lample, G., Barrault, L., and Baroni, M. (2018) · 2018
Cited alongside, same era.
Adversarial removal of demographic attributes from text data
Elazar, Y., and Goldberg, Y. (2018) · 2018
Cited alongside, same era.
Comparing deep learning and concept extraction based methods for patient phenotyping from clinical narratives
Gehrmann, S., Dernoncourt, F., Li, Y., Carlson, E. T., Wu, J. T., Welt, J., Foote Jr, J., Moseley, E. T., Grant, D. W., Tyler, P. D., et al. (2018) · 2018
Cited alongside, same era.
Under the hood: Using diagnostic classifiers to investigate and improve how language models track agreement information
Giulianelli, M., Harding, J., Mohnert, F., Hupkes, D., and Zuidema, W. (2018) · 2018
Cited alongside, same era.
Visualisation and ‘diagnostic classifiers’ reveal how recurrent and recursive neural networks process hierarchical structure
Hupkes, D., Veldhoen, S., and Zuidema, W. (2018) · 2018
Quantifying social biases in contextual word representations
Kurita, K., Vyas, N., Pareek, A., Black, A. W., and Tsvetkov, Y. (2019) · 2019
Later among the works it cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019) · 2019
Later among the works it cites.
Fairness through causal awareness: Learning causal latent-variable models for biased data
Madras, D., Creager, E., Pitassi, T., and Zemel, R. (2019) · 2019
Later among the works it cites.
Layer-wise relevance propagation: an overview
Montavon, G., Binder, A., Lapuschkin, S., Samek, W., and Müller, K.-R. (2019) · 2019
Later among the works it cites.
Fast parallel algorithms for feature selection.
Qian, S., and Singer, Y. (2019) · 2019
Later among the works it cites.
Reducing gender bias in word-level language models with a gender-equalizing loss function
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Examining gender and race bias in two hundred sentiment analysis systems
Kiritchenko, S., and Mohammad, S. (2018) · 2018
Cited alongside, same era.
Beyond word importance: Contextual decomposition to extract interactions from LSTMs
Murdoch, W. J., Liu, P. J., and Yu, B. (2018) · 2018
Cited alongside, same era.
Stress test evaluation for natural language inference
Naik, A., Ravichander, A., Sadeh, N., Rose, C., and Neubig, G. (2018) · 2018
Cited alongside, same era.
Gender bias in coreference resolution
Rudinger, R., Naradowsky, J., Leonard, B., and Van Durme, B. (2018) · 2018
Cited alongside, same era.
Mind the gap: A balanced corpus of gendered ambiguous pronouns
Webster, K., Recasens, M., Axelrod, V., and Baldridge, J. (2018) · 2018
Cited alongside, same era.
Gender bias in coreference resolution: Evaluation and debiasing methods
Zhao, J., Wang, T., Yatskar, M., Ordonez, V., and Chang, K.-W. (2018a) · 2018
Cited alongside, same era.
Learning gender-neutral word embeddings
Zhao, J., Zhou, Y., Li, Z., Wang, W., and Chang, K.-W. (2018b) · 2018
Cited alongside, same era.
Qian, Y., Muaz, U., Zhang, B., and Hyun, J. W. (2019) · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. (2019) · 2019
Later among the works it cites.
What’s in a name? Reducing bias in bios without access to protected attributes
Romanov, A., De-Arteaga, M., Wallach, H., Chayes, J., Borgs, C., Chouldechova, A., Geyik, S., Kenthapadi, K., Rumshisky, A., and Kalai, A. (2019) · 2019
Later among the works it cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T. (2019) · 2019
Later among the works it cites.
Effective sentence scoring method using bert for speech recognition
Shin, J., Lee, Y., and Jung, K. (2019) · 2019
Later among the works it cites.
Assessing social and intersectional biases in contextualized word representations
Tan, Y. C., and Celis, L. E. (2019) · 2019
Later among the works it cites.
BERT rediscovers the classical NLP pipeline
Tenney, I., Das, D., and Pavlick, E. (2019) · 2019
Later among the works it cites.
A multiscale visualization of attention in the transformer model
Vig, J. (2019) · 2019
Later among the works it cites.
Analyzing the structure of attention in a transformer language model
Vig, J., and Belinkov, Y. (2019) · 2019
Later among the works it cites.
BERT has a mouth, and it must speak: BERT as a Markov random field language model
Wang, A., and Cho, K. (2019) · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R. R., and Le, Q. V. (2019) · 2019
Later among the works it cites.
Gender bias in contextualized word embeddings
Zhao, J., Wang, T., Yatskar, M., Cotterell, R., Ordonez, V., and Chang, K.-W. (2019a) · 2019
Later among the works it cites.
Gender bias in contextualized word embeddings
Zhao, J., Wang, T., Yatskar, M., Cotterell, R., Ordonez, V., and Chang, K.-W. (2019b) · 2019
Later among the works it cites.
Causal interpretations of black-box models
Zhao, Q., and Hastie, T. (2019) · 2019
Later among the works it cites.
Toward gender-inclusive coreference resolution
Cao, Y. T., and Daumé III, H. (2020) · 2020
Closest in time.
When bert forgets how to POS: Amnesic probing of linguistic properties and MLM predictions
Elazar, Y., Ravfogel, S., Jacovi, A., and Goldberg, Y. (2020) · 2020
Closest in time.
Causalm: Causal model explanation through counterfactual language models
Feder, A., Oved, N., Shalit, U., and Reichart, R. (2020) · 2020
Closest in time.
Reducing sentiment bias in language models via counterfactual evaluation
Huang, P.-S., Zhang, H., Jiang, R., Stanforth, R., Welbl, J., Rae, J., Maini, V., Yogatama, D., and Kohli, P. (2020) · 2020
Closest in time.
Gender Bias in Neural Natural Language Processing
Lu, K., Mardziel, P., Wu, F., Amancharla, P., and Datta, A. (2020) · 2020
Closest in time.
Masked language model scoring
Salazar, J., Liang, D., Nguyen, T. Q., and Kirchhoff, K. (2020) · 2020
Closest in time.
Investigating gender bias in language models using causal mediation analysis
Vig, J., Gehrmann, S., Belinkov, Y., Qian, S., Nevo, D., Singer, Y., and Shieber, S. (2020) · 2020
Closest in time.
A causal inference method for reducing gender bias in word embedding relations
Yang, Z., and Feng, J. (2020) · 2020
Closest in time.