Fetching the paper…
Reading the bibliography…
One well motivated explanation method for classifiers leverages counterfactuals which are hypothetical events identical to real observations in all aspects except for one feature.
Causality
Judea Pearl · 2009
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai · 2016
Earlier work this paper cites.
Counterfactual fairness
Matt J Kusner, Joshua Loftus, Chris Russell, and Ricardo Silva · 2017
Earlier work this paper cites.
Elements of causal inference: foundations and learning algorithms
Jonas Peters, Dominik Janzing, and Bernhard Schölkopf · 2017
Earlier work this paper cites.
Controllable invariance through adversarial feature learning
Qizhe Xie, Zihang Dai, Yulun Du, Eduard Hovy, and Graham Neubig · 2017
Earlier work this paper cites.
Adversarial removal of demographic attributes from text data
Yanai Elazar and Yoav Goldberg · 2018
Earlier work this paper cites.
Bias in bios: A case study of semantic representation bias in a high-stakes setting
Maria De-Arteaga, Alexey Romanov, Hanna Wallach, Jennifer Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Geyik, Krishnaram Kenthapadi, and Adam Tauman Kalai · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Counterfactual fairness in text classification through robustness
Sahaj Garg, Vincent Perot, Nicole Limtiaco, Ankur Taly, Ed H Chi, and Alex Beutel · 2019
Earlier work this paper cites.
Learning the difference that makes a difference with counterfactually-augmented data
Divyansh Kaushik, Eduard Hovy, and Zachary Lipton · 2020
Cited alongside, same era.
Explaining machine learning classifiers through diverse counterfactual explanations
Ramaravind K Mothilal, Amit Sharma, and Chenhao Tan · 2020
Cited alongside, same era.
Null it out: Guarding protected attributes by iterative nullspace projection
Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg · 2020
Cited alongside, same era.
Amnesic probing: Behavioral explanation with amnesic counterfactuals
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg · 2021
Cited alongside, same era.
CausaLM: Causal model explanation through counterfactual language models
Amir Feder, Nadav Oved, Uri Shalit, and Roi Reichart · 2021
Cited alongside, same era.
Explaining the efficacy of counterfactually augmented data
Counterfactual generator: A weakly-supervised method for named entity recognition
Xiangji Zeng, Yunliang Li, Yuchen Zhai, and Yin Zhang · 2021
Later among the works it cites.
CEBaB: Estimating the causal effects of real-world concepts on nlp model behavior
Eldar D Abraham, Karel D’Oosterlinck, Amir Feder, Yair Gat, Atticus Geiger, Christopher Potts, Roi Reichart, and Zhengxuan Wu · 2022
Later among the works it cites.
DoCoGen: Domain counterfactual generation for low resource domain adaptation
Nitay Calderon, Eyal Ben-David, Amir Feder, and Roi Reichart · 2022
Later among the works it cites.
Adversarial concept erasure in kernel space
Shauli Ravfogel, Francisco Vargas, Yoav Goldberg, and Ryan Cotterell · 2022
Later among the works it cites.
Robust concept erasure via kernelized rate-distortion maximization
Somnath Basu Roy Chowdhury, Nicholas Monath, Kumar Avinava Dubey, Amr Ahmed, and Snigdha Chaturvedi · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Divyansh Kaushik, Amrith Setlur, Eduard H Hovy, and Zachary Chase Lipton · 2021
Cited alongside, same era.
Generate your counterfactuals: Towards controlled counterfactual generation for text
Nishtha Madaan, Inkit Padhi, Naveen Panwar, and Diptikalyan Saha · 2021
Cited alongside, same era.
Counterfactual invariance to spurious correlations in text classification
Victor Veitch, Alexander D’Amour, Steve Yadlowsky, and Jacob Eisenstein · 2021
Cited alongside, same era.
Nora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell, Edward Raff, and Stella Biderman · 2023
Later among the works it cites.
Log-linear guardedness and its implications
Shauli Ravfogel, Yoav Goldberg, and Ryan Cotterell · 2023
Later among the works it cites.
Gold doesn’t always glitter: Spectral removal of linear and nonlinear guarded attribute information
Shun Shao, Yftah Ziser, and Shay B. Cohen · 2023
Later among the works it cites.