Fetching the paper…
Reading the bibliography…
Recent years have witnessed an increasing number of interpretation methods being developed for improving transparency of NLP models.
Evaluating explanation without ground truth in interpretable machine learning
Fan Yang, Mengnan Du, and Xia Hu. 2019 · 1907
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Towards faithfully interpretable nlp systems: How should we define and evaluate faithfulness?
Alon Jacovi and Yoav Goldberg. 2020 · 2004
Earlier work this paper cites.
Probability and statistical inference
Robert V Hogg, Elliot A Tanis, and Dale L Zimmerman. 2010 · 2010
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2013 · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Extraction of salient sentences from labelled documents
Misha Denil, Alban Demiraj, and Nando De Freitas. 2014 · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
A compositional and interpretable semantic space
Alona Fyshe, Leila Wehbe, Partha Talukdar, Brian Murphy, and Tom Mitchell. 2015 · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
Interpretable deep models for icu outcome prediction
Zhengping Che, Sanjay Purushotham, Robinder Khemani, and Yan Liu. 2016 · 2016
Earlier work this paper cites.
Interpretation of prediction models using the input gradient
Yotam Hechtlinger. 2016 · 2016
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Thomas N Kipf and Max Welling. 2016 · 2016
Earlier work this paper cites.
Understanding neural networks through representation erasure
Jiwei Li, Will Monroe, and Dan Jurafsky. 2016 · 2016
Cited alongside, same era.
Why should i trust you?: Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Cited alongside, same era.
Interpretable explanations of black boxes by meaningful perturbation
Ruth C Fong and Andrea Vedaldi. 2017 · 2017
Cited alongside, same era.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang. 2017 · 2017
Cited alongside, same era.
Dynamic routing between capsules
Sara Sabour, Nicholas Frosst, and Geoffrey E Hinton. 2017 · 2017
Cited alongside, same era.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Explanations can be manipulated and geometry is to blame
Ann-Kathrin Dombrowski, Maximillian Alber, Christopher Anders, Marcel Ackermann, Klaus-Robert Müller, and Pan Kessel. 2019 · 2019
Later among the works it cites.
Interpretation of neural networks is fragile
Amirata Ghorbani, Abubakar Abid, and James Zou. 2019 · 2019
Later among the works it cites.
Towards a deep and unified understanding of deep neural models in nlp
Chaoyu Guan, Xiting Wang, Quanshi Zhang, Runjin Chen, Di He, and Xing Xie. 2019 · 2019
Later among the works it cites.
Applying deep learning to airbnb search
Malay Haldar, Mustafa Abdool, Prashant Ramanathan, Tao Xu, Shulin Yang, Huizhong Duan, Qing Zhang, Nick Barrow-Williams, Bradley C Turnbull, Brendan M Collins, et al. 2019 · 2019
Later among the works it cites.
Attention is not explanation
Sarthak Jain and Byron C Wallace. 2019 · 2019
Later among the works it cites.
The effects of meaningful and meaningless explanations on trust and perceived system accuracy in intelligent systems
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. 2017 · 2017
Cited alongside, same era.
Smoothgrad: removing noise by adding noise
Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda Viégas, and Martin Wattenberg. 2017 · 2017
Cited alongside, same era.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio. 2017 · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Towards explanation of dnn-based prediction with guided feature inversion
Mengnan Du, Ninghao Liu, Qingquan Song, and Xia Hu. 2018 · 2018
Cited alongside, same era.
Mahsan Nourani, Samia Kabir, Sina Mohseni, and Eric D Ragan. 2019 · 2019
Later among the works it cites.
Word2sense: Sparse interpretable word embeddings
Abhishek Panigrahi, Harsha Vardhan Simhadri, and Chiranjib Bhattacharyya. 2019 · 2019
Later among the works it cites.
Human-centered artificial intelligence and machine learning
Mark O Riedl. 2019 · 2019
Later among the works it cites.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Cynthia Rudin. 2019 · 2019
Later among the works it cites.
Is attention interpretable?
Sofia Serrano and Noah A Smith. 2019 · 2019
Later among the works it cites.
Allennlp interpret: A framework for explaining predictions of nlp models
Eric Wallace, Jens Tuyls, Junlin Wang, Sanjay Subramanian, Matt Gardner, and Sameer Singh. 2019 · 2019
Later among the works it cites.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter. 2019 · 2019
Later among the works it cites.
Theoretically principled trade-off between robustness and accuracy
Hongyang Zhang, Yaodong Yu, Jiantao Jiao, Eric Xing, Laurent El Ghaoui, and Michael Jordan. 2019 · 2019
Later among the works it cites.
The polar framework: Polar opposites enable interpretability of pre-trained word embeddings
Binny Mathew, Sandipan Sikdar, Florian Lemmerich, and Markus Strohmaier. 2020 · 2020
Closest in time.