Fetching the paper…
Reading the bibliography…
A principle behind dozens of attribution methods is to take the prediction difference between before-and-after an input feature (here, a token) is removed as its attribution.
" why should you trust my explanation?" understanding uncertainty in lime explanations
Yujia Zhang, Kuangyan Song, Yiming Sun, Sarah Tan, and Madeleine Udell. 2019 · 1904
Earlier work this paper cites.
Explaining classifiers with causal concept effect (cace)
Yash Goyal, Amir Feder, Uri Shalit, and Been Kim. 2019 · 1907
Earlier work this paper cites.
Explaining classifications for individual instances
Marko Robnik-Šikonja and Igor Kononenko. 2008 · 2008
Earlier work this paper cites.
Causality
Judea Pearl. 2009 · 2009
Earlier work this paper cites.
Feature removal is a unifying principle for model explanation methods
Ian Covert, Scott Lundberg, and Su-In Lee. 2020 · 2011
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013a · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013b · 2013
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D Zeiler and Rob Fergus. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Understanding neural networks through representation erasure
Jiwei Li, Will Monroe, and Dan Jurafsky. 2016 · 2016
Earlier work this paper cites.
Why should i trust you?: Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Learning deep features for discriminative localization
Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba. 2016 · 2016
Cited alongside, same era.
Explaining recurrent neural network predictions in sentiment analysis
Leila Arras, Grégoire Montavon, Klaus-Robert Müller, and Wojciech Samek. 2017 · 2017
Cited alongside, same era.
Real time image saliency for black box classifiers
Piotr Dabkowski and Yarin Gal. 2017 · 2017
Cited alongside, same era.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee. 2017 · 2017
Cited alongside, same era.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim. 2018 · 2018
Cited alongside, same era.
e-snli: Natural language inference with natural language explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018 · 2018
Incorporating priors with feature attribution on text classification
Frederick Liu and Besim Avci. 2019 · 2019
Later among the works it cites.
Explaining image classifiers by removing input features using generative models
Chirag Agarwal and Anh Nguyen. 2020 · 2020
Later among the works it cites.
Sam: The sensitivity of attribution methods to hyperparameters
Naman Bansal, Chirag Agarwal, and Anh Nguyen. 2020 · 2020
Later among the works it cites.
ERASER: A benchmark to evaluate rationalized NLP models
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace. 2020 · 2020
Later among the works it cites.
Considering likelihood in NLP classification explanations with occlusion and language modeling
David Harbecke and Christoph Alt. 2020 · 2020
Later among the works it cites.
Pretrained models — transformers 3.3.0 documentation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Looking beyond the surface: A challenge set for reading comprehension over multiple sentences
Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. 2018 · 2018
Cited alongside, same era.
Explaining image classifiers by counterfactual generation
Chun-Hao Chang, Elliot Creager, Anna Goldenberg, and David Duvenaud. 2019 · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Understanding deep networks via extremal perturbations and smooth masks
Ruth Fong, Mandela Patrick, and Andrea Vedaldi. 2019 · 2019
Cited alongside, same era.
A benchmark for interpretability methods in deep neural networks
Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim. 2019 · 2019
Cited alongside, same era.
On human predictions with explanations and predictions of machine learning models: A case study on deception detection
Vivian Lai and Chenhao Tan. 2019 · 2019
Cited alongside, same era.
Huggingface. 2020 · 2020
Later among the works it cites.
Towards hierarchical importance attribution: Explaining compositional semantics for neural sequence models
Xisen Jin, Zhongyu Wei, Junyi Du, Xiangyang Xue, and Xiang Ren. 2020 · 2020
Later among the works it cites.
Interpretation of NLP models through input marginalization
Siwon Kim, Jihun Yi, Eunji Kim, and Sungroh Yoon. 2020 · 2020
Later among the works it cites.
The out-of-distribution problem in explainability and search methods for feature importance explanations
Peter Hase, Harry Xie, and Mohit Bansal. 2021 · 2021
Closest in time.
marcotcr/lime: Lime: Explaining the predictions of any machine learning classifier
Ribeiro. 2021 · 2021
Closest in time.
Teach me to explain: A review of datasets for explainable nlp
Sarah Wiegreffe and Ana Marasović. 2021 · 2021
Closest in time.