Approximate causal abstractions
Beckers, S., Eberhardt, F., and Halpern, J. Y · 2020
Later among the works it cites.
Counterfactuals uncover the modular structure of deep generative models
Besserve, M., Mehrjou, A., Sun, R., and Schölkopf, B · 2020
Later among the works it cites.
Amnesic probing: Behavioral explanation with amnesic counterfactuals
Elazar, Y., Ravfogel, S., Jacovi, A., and Goldberg, Y · 2020
Later among the works it cites.
Neural natural language inference models partially embed theories of lexical entailment and negation
Geiger, A., Richardson, K., and Potts, C · 2020
Later among the works it cites.
Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?
Jacovi, A. and Goldberg, Y · 2020
Later among the works it cites.
Null it out: Guarding protected attributes by iterative nullspace projection
Ravfogel, S., Elazar, Y., Gonen, H., Twiton, M., and Goldberg, Y · 2020
Later among the works it cites.
Probing the probing paradigm: Does probing accuracy entail task relevance?, 2020
Ravichander, A., Belinkov, Y., and Hovy, E · 2020
Later among the works it cites.
A benchmark for systematic generalization in grounded language understanding
Ruis, L., Andreas, J., Baroni, M., Bouchacourt, D., and Lake, B. M · 2020
Later among the works it cites.
Discovering the compositional structure of vector representations with role learning networks
Soulos, P., McCoy, R. T., Linzen, T., and Smolensky, P · 2020
Later among the works it cites.
Causal mediation analysis for interpreting neural nlp: The case of gender bias, 2020
Vig, J., Gehrmann, S., Belinkov, Y., Qian, S., Nevo, D., Singer, Y., and Shieber, S · 2020
Later among the works it cites.
Are neural nets modular? inspecting functional modularity through differentiable weight masks
Csordás, R., van Steenkiste, S., and Schmidhuber, J · 2021
Closest in time.
Improving performance of deep learning models with axiomatic attribution priors and expected gradients
Erion, G., Janizek, J. D., Sturmfels, P., Lundberg, S. M., and Lee, S.-I · 2021
Closest in time.
Causal abstractions of neural networks
Original
Geiger, A., Lu, H., Icard, T., and Potts, C · 2021
Closest in time.
Counterfactual data augmentation for neural machine translation
Liu, Q., Kusner, M., and Blunsom, P · 2021
Closest in time.
Causal effects of linguistic properties
Pryzant, R., Card, D., Jurafsky, D., Veitch, V., and Sridhar, D · 2021
Closest in time.
ReaSCAN: Compositional reasoning in language grounding
Original
Wu, Z., Kreiss, E., Ong, D. C., and Potts, C · 2021
Closest in time.
Pointer value retrieval: A new benchmark for understanding the limits of neural network generalization
Original
Zhang, C., Raghu, M., Kleinberg, J. M., and Bengio, S · 2021
Closest in time.
Locating and editing factual associations in gpt, 2022
Original
Meng, K., Bau, D., Andonian, A., and Belinkov, Y · 2022
Closest in time.