Fetching the paper…
Reading the bibliography…
Gradient-based analysis methods, such as saliency map visualizations and adversarial input perturbations, have found widespread use in interpreting neural NLP models due to their simplicity, flexibility, and most importantly, their faithfulness.
NLTK: the natural language toolkit
Edward Loper and Steven Bird. 2002 · 2002
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning. 2015 · 2015
Earlier work this paper cites.
Layer normalization
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Why should I trust you?: Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim. 2017 · 2017
Earlier work this paper cites.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang. 2017 · 2017
Earlier work this paper cites.
The (un)reliability of saliency methods
Pieter-Jan Kindermans, Sara Hooker, Julius Adebayo, Maximilian Alber, Kristof T. Schütt, Sven Dähne, Dumitru Erhan, and Been Kim. 2017 · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott Lundberg and Su-In Lee. 2017 · 2017
Earlier work this paper cites.
Right for the right reasons: Training differentiable models by constraining their explanations
Andrew Slavin Ross, Michael C Hughes, and Finale Doshi-Velez. 2017 · 2017
Earlier work this paper cites.
SmoothGrad: removing noise by adding noise
Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda B. Viégas, and Martin Wattenberg. 2017 · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim. 2018 · 2018
Cited alongside, same era.
Do explanations make VQA models more predictable to a human?
Arjun Chandrasekaran, Viraj Prabhu, Deshraj Yadav, Prithvijit Chattopadhyay, and Devi Parikh. 2018 · 2018
Cited alongside, same era.
HotFlip: White-box adversarial examples for text classification
Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou. 2018 · 2018
Cited alongside, same era.
Pathologies of neural models make interpretations difficult
Shi Feng, Eric Wallace, Alvin Grissom II, Mohit Iyyer, Pedro Rodriguez, and Jordan Boyd-Graber. 2018 · 2018
Cited alongside, same era.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel R. Bowman, and Noah A. Smith. 2018 · 2018
Interpretation of neural networks is fragile
Amirata Ghorbani, Abubakar Abid, and James Zou. 2019 · 2019
Later among the works it cites.
Fooling neural network interpretations via adversarial model manipulation
Juyeon Heo, Sunghwan Joo, and Taesup Moon. 2019 · 2019
Later among the works it cites.
Attention is not explanation
Sarthak Jain and Byron C Wallace. 2019 · 2019
Later among the works it cites.
On human predictions with explanations and predictions of machine learning models: A case study on deception detection
Vivian Lai and Chenhao Tan. 2019 · 2019
Later among the works it cites.
Incorporating priors with feature attribution on text classification
Frederick Liu and Besim Avci. 2019 · 2019
Later among the works it cites.
Is attention interpretable?
Sofia Serrano and Noah A Smith. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Beyond word importance: Contextual decomposition to extract interactions from LSTMs
W. James Murdoch, Peter J. Liu, and Bin Yu. 2018 · 2018
Cited alongside, same era.
Anchors: High-precision model-agnostic explanations
Marco Túlio Ribeiro, Sameer Singh, and Carlos Guestrin. 2018 · 2018
Cited alongside, same era.
Improving the adversarial robustness and interpretability of deep neural networks by regularizing their input gradients
Andrew Slavin Ross and Finale Doshi-Velez. 2018 · 2018
Cited alongside, same era.
Bias in bios
Maria De-Arteaga, Alexey Romanov, Hanna Wallach, Jennifer Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Geyik, Krishnaram Kenthapadi, and Adam Tauman Kalai. 2019 · 2019
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Explanations can be manipulated and geometry is to blame
Ann-Kathrin Dombrowski, Maximillian Alber, Christopher Anders, Marcel Ackermann, Klaus-Robert Müller, and Pan Kessel. 2019 · 2019
Cited alongside, same era.
Fooling network interpretation in image classification
Akshayvarun Subramanya, Vipin Pillai, and Hamed Pirsiavash. 2019 · 2019
Later among the works it cites.
How to manipulate CNNs to make them lie: the GradCAM case
Tom Viering, Ziqi Wang, Marco Loog, and Elmar Eisemann. 2019 · 2019
Later among the works it cites.
AllenNLP Interpret: A framework for explaining predictions of NLP models
Eric Wallace, Jens Tuyls, Junlin Wang, Sanjay Subramanian, Matt Gardner, and Sameer Singh. 2019 · 2019
Later among the works it cites.
You shouldn’t trust me: Learning models which conceal unfairness from multiple explanation methods
Botty Dimanov, Umang Bhatt, Mateja Jamnik, and Adrian Weller. 2020 · 2020
Closest in time.
“How do I fool you?” manipulating user trust via misleading black box explanations
Himabindu Lakkaraju and Osbert Bastani. 2020 · 2020
Closest in time.
Learning to deceive with attention-based explanations
Danish Pruthi, Mansi Gupta, Bhuwan Dhingra, Graham Neubig, and Zachary C Lipton. 2020 · 2020
Closest in time.
Interpretations are useful: penalizing explanations to align neural networks with prior knowledge
Laura Rieger, Chandan Singh, W James Murdoch, and Bin Yu. 2020 · 2020
Closest in time.
Fooling LIME and SHAP: Adversarial attacks on post hoc explanation methods
Dylan Slack, Sophie Hilgard, Emily Jia, Sameer Singh, and Himabindu Lakkaraju. 2020 · 2020
Closest in time.