Fetching the paper…
Reading the bibliography…
An increasing number of machine learning models have been deployed in domains with high stakes such as finance and healthcare.
Why should you trust my interpretation? understanding uncertainty in lime predictions
Hui Fen Tan, Kuangyan Song, Madeilene Udell, Yiming Sun, and Yujia Zhang. 2019 · 1904
Earlier work this paper cites.
Did the model understand the question?
Pramod Kaushik Mudrakarta, Ankur Taly, Mukund Sundararajan, and Kedar Dhamdhere. 2018 · 1906
Earlier work this paper cites.
Sofia Serrano and Noah A. Smith. 2019 · 1906
Earlier work this paper cites.
Muhammad Rehman Zafar and Naimul Mefraz Khan. 2019 · 1906
Earlier work this paper cites.
Attention interpretability across nlp tasks
Shikhar Vashishth, Shyam Upadhyay, Gaurav Singh Tomar, and Manaal Faruqui. 2019 · 1909
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam M. Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 1910
Earlier work this paper cites.
A value fo n-person games
Lloyd Shapley. 1953 · 1953
Earlier work this paper cites.
Toward interpretability of dual-encoder models for dialogue response suggestions
Yitong Li, Dianqi Li, Sushant Prakash, and Peng Wang. 2020 · 2003
Earlier work this paper cites.
Visualizing data using t-sne
Laurens van der Maaten and Geoffrey E. Hinton. 2008 · 2008
Earlier work this paper cites.
Shap values for explaining cnn-based text classification models
Wei Zhao, Tarun Joshi, Vijayan N. Nair, and A. Sudjianto. 2020 · 2008
Earlier work this paper cites.
The magical mystery four: How is working memory capacity limited, and why?
Nelson Cowan. 2010 · 2010
Earlier work this paper cites.
Gradient-based analysis of nlp models is manipulable
Junlin Wang, Jens Tuyls, Eric Wallace, and Sameer Singh. 2020 · 2010
Earlier work this paper cites.
Deconvolutional networks
Matthew D. Zeiler, Dilip Krishnan, Graham W. Taylor, and Rob Fergus. 2010 · 2010
Earlier work this paper cites.
A survey on neural network interpretability
Yu Zhang, Peter Tiño, Ales Leonardis, and Ke Tang. 2020 · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Gregory S. Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Distributional vectors encode referential attributes
Abhijeet Gupta, Gemma Boleda, Marco Baroni, and Sebastian Padó. 2015 · 2015
Earlier work this paper cites.
How well do distributional models capture different types of semantic knowledge?
Dana Rubinstein, Effi Levi, Roy Schwartz, and Ari Rappoport. 2015 · 2015
Earlier work this paper cites.
Striving for simplicity: The all convolutional net
Jost Tobias Springenberg, Alexey Dosovitskiy, Thomas Brox, and Martin A. Riedmiller. 2015 · 2015
Earlier work this paper cites.
Explaining predictions of non-linear classifiers in NLP
Leila Arras, Franziska Horn, Grégoire Montavon, Klaus-Robert Müller, and Wojciech Samek. 2016 · 2016
Earlier work this paper cites.
Layer-wise relevance propagation for neural networks with local renormalization layers
Alexander Binder, Grégoire Montavon, Sebastian Lapuschkin, Klaus-Robert Müller, and Wojciech Samek. 2016 · 2016
Earlier work this paper cites.
Word embedding evaluation and combination
Sahar Ghannay, Benoit Favre, Yannick Estève, and Nathalie Camelin. 2016 · 2016
Earlier work this paper cites.
Evaluating Embeddings using Syntax-based Classification Tasks as a Proxy for Parser Performance
Arne Köhn. 2016 · 2016
Earlier work this paper cites.
Visualizing and understanding neural models in NLP
Jiwei Li, Xinlei Chen, Eduard Hovy, and Dan Jurafsky. 2016 · 2016
Earlier work this paper cites.
The mythos of model interpretability
Zachary Chase Lipton. 2016 · 2016
Earlier work this paper cites.
Hierarchical question-image co-attention for visual question answering
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. 2016 · 2016
Earlier work this paper cites.
Does string-based neural MT learn source syntax?
Xing Shi, Inkit Padhi, and Kevin Knight. 2016 · 2016
Earlier work this paper cites.
Fine-grained analysis of sentence embeddings using auxiliary prediction tasks
Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg. 2017 · 2017
Earlier work this paper cites.
What do neural machine translation models learn about morphology?
Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, and James Glass. 2017 · 2017
Earlier work this paper cites.
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017 · 2017
Earlier work this paper cites.
Interpretability of deep learning models: A survey of results
Supriyo Chakraborty, Richard Tomsett, Ramya Raghavendra, Daniel Harborne, Moustafa Alzantot, Federico Cerutti, Mani Srivastava, Alun Preece, Simon Julier, Raghuveer M. Rao, Troy D. Kelley, Dave Braines, Murat Sensoy, Christopher J. Willis, and Prudhvi Gurram. 2017 · 2017
Earlier work this paper cites.
Hotflip: White-box adversarial examples for NLP
Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou. 2017 · 2017
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom. 2017 · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M. Lundberg and Su-In Lee. 2017 · 2017
Earlier work this paper cites.
Rotated word vector representations and their interpretability
Sungjoon Park, JinYeong Bak, and Alice H. Oh. 2017 · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje. 2017 · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Cited alongside, same era.
Attention Is All You Need
Ashish Vaswani. 2017 · 2017
Cited alongside, same era.
On the robustness of interpretability methods
David Alvarez-Melis and T. Jaakkola. 2018 · 2018
Cited alongside, same era.
Understanding the representational power of neural retrieval models using NLP tasks
Daniel Cohen, Brendan O’Connor, and W. Bruce Croft. 2018 · 2018
Cited alongside, same era.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018 · 2018
Cited alongside, same era.
HotFlip : White-Box Adversarial Examples for Text Classification
Generating Counterfactual and Contrastive Explanations using SHAP
Shubham Rathi. 2019 · 2019
Later among the works it cites.
What do you learn from context ?
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R. Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R. Bowman, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Later among the works it cites.
A multiscale visualization of attention in the transformer model
Jesse Vig. 2019 · 2019
Later among the works it cites.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter. 2019 · 2019
Later among the works it cites.
Language Modeling Teaches You More than Translation Does: Lessons Learned Through Auxiliary Syntactic Task Analysis
Kelly Zhang and Samuel Bowman. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou. 2018 · 2018
Cited alongside, same era.
Adversarial attacks against medical deep learning systems
Samuel G. Finlayson, Isaac S. Kohane, and Andrew L. Beam. 2018 · 2018
Cited alongside, same era.
Under the hood: Using diagnostic classifiers to investigate and improve how language models track agreement information
Mario Giulianelli, Jack Harding, Florian Mohnert, Dieuwke Hupkes, and Willem Zuidema. 2018 · 2018
Cited alongside, same era.
Milo Honegger. 2018 · 2018
Cited alongside, same era.
Visualisation and ’diagnostic classifiers’ reveal how recurrent and recursive neural networks process hierarchical structure
Dieuwke Hupkes and Willem Zuidema. 2018 · 2018
Cited alongside, same era.
Consistent individualized feature attribution for tree ensembles
Scott M. Lundberg, Gabriel G. Erion, and Su-In Lee. 2018 · 2018
Cited alongside, same era.
Beyond polarity: Interpretable financial sentiment analysis with hierarchical query-driven attention
Ling Luo, Xiang Ao, Feiyang Pan, Jin Wang, Tong Zhao, Ningzi Yu, and Qing He. 2018 · 2018
Cited alongside, same era.
Kieran Browne and Ben Swift. 2020 · 2020
Later among the works it cites.
Evaluating models’ local decision boundaries via contrast sets
Matt Gardner, Yoav Artzi, Victoria Basmova, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, Nitish Gupta, Hanna Hajishirzi, Gabriel Ilharco, Daniel Khashabi, Kevin Lin, Jiangming Liu, Nelson F. Liu, Phoebe Mulcaire, Qiang Ning, Sameer Singh, Noah A. Smith, Sanjay Subramanian, Reut Tsarfaty, Eric Wallace, Ally Zhang, and Ben Zhou. 2020 · 2020
Later among the works it cites.
Explain2Attack: Text adversarial attacks via cross-domain interpretability
Mahmoud Hossam, Trung Le, He Zhao, and Dinh Phung. 2020 · 2020
Later among the works it cites.
Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?
Alon Jacovi and Yoav Goldberg. 2020 · 2020
Later among the works it cites.
NILE : Natural Language Inference with Faithful Natural Language Explanations
Sawan Kumar and Partha Talukdar. 2020 · 2020
Later among the works it cites.
Are multilingual neural machine translation models better at capturing linguistic features?
David Mareek, Hande Celikkanat, Miikka Silfverberg, Vinit Ravishankar, and Jrg Tiedemann. 2020 · 2020
Later among the works it cites.
Causal Interpretability for Machine Learning - Problems, Methods and Evaluation
Raha Moraffah, Mansooreh Karami, Ruocheng Guo, Adrienne Raglin, and Huan Liu. 2020 · 2020
Later among the works it cites.
What does BERT know about books, movies and music? Probing BERT for Conversational Recommendation
Gustavo Penha and Claudia Hauff. 2020 · 2020
Later among the works it cites.
Information-theoretic probing for linguistic structure
Tiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod, Adina Williams, and Ryan Cotterell. 2020 · 2020
Later among the works it cites.
Learning to deceive with attention-based explanations
Danish Pruthi, Mansi Gupta, Bhuwan Dhingra, Graham Neubig, and Zachary Chase Lipton. 2020 · 2020
Later among the works it cites.
A primer in bertology: What we know about how bert works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020 · 2020
Later among the works it cites.
LINSPECTOR: Multilingual probing tasks for word representations
Gözde Gül Şahin, Clara Vania, Ilia Kuznetsov, and Iryna Gurevych. 2020 · 2020
Later among the works it cites.
Fooling lime and shap: Adversarial attacks on post hoc explanation methods
Dylan Slack, Sophie Hilgard, Emily Jia, Sameer Singh, and Himabindu Lakkaraju. 2020 · 2020
Later among the works it cites.
Probing for referential information in language models
Ionut-Teodor Sorodoc, Kristina Gulordava, and Gemma Boleda. 2020 · 2020
Later among the works it cites.
Select, answer and explain: Interpretable multi-hop reading comprehension over multiple documents
Ming Tu, Kevin Huang, Guangtao Wang, Jing Huang, Xiaodong He, and Bowen Zhou. 2020 · 2020
Later among the works it cites.
Exploring What Is Encoded in Distributional Word Vectors: A Neurobiologically Motivated Analysis
Akira Utsumi. 2020 · 2020
Later among the works it cites.
Causal Mediation Analysis for Interpreting Neural NLP: The Case of Gender Bias
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Simas Sakenis, Jason Huang, Yaron Singer, and Stuart Shieber. 2020 · 2020
Later among the works it cites.
Information-theoretic probing with minimum description length
Elena Voita and Ivan Titov. 2020 · 2020
Later among the works it cites.
Assessing phrasal representation and composition in transformers
Lang Yu and Allyson Ettinger. 2020 · 2020
Later among the works it cites.
Fake news spreaders profiling using n-grams of various types and shap-based feature selection
Fazlourrahman Balouchzahi, Grigori Sidorov, and Hosahalli Lakshmaiah Shashirekha. 2021 · 2021
Later among the works it cites.
Probing Classifiers: Promises, Shortcomings, and Advances
Yonatan Belinkov. 2021 · 2021
Later among the works it cites.
Causal analysis of syntactic agreement mechanisms in neural language models
Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart Shieber, Tal Linzen, and Yonatan Belinkov. 2021 · 2021
Later among the works it cites.
The Intriguing Relation Between Counterfactual Explanations and Adversarial Examples
Timo Freiesleben. 2021 · 2021
Later among the works it cites.
Contrastive Explanations for Model Interpretability
Alon Jacovi, Swabha Swayamdipta, Shauli Ravfogel, Yanai Elazar, Yejin Choi, and Yoav Goldberg. 2021 · 2021
Later among the works it cites.
Post-hoc Interpretability for Neural NLP: A Survey
Andreas Madsen, Siva Reddy, and Sarath Chandar. 2021 · 2021
Later among the works it cites.
Probing the probing paradigm: Does probing accuracy entail task relevance?
Abhilasha Ravichander, Yonatan Belinkov, and Eduard Hovy. 2021 · 2021
Later among the works it cites.
Does BERT Understand Idioms? A Probing-Based Empirical Study of BERT Encodings of Idioms
Minghuan Tan and Jing Jiang. 2021 · 2021
Later among the works it cites.
Understanding Pure Character-Based Neural Machine Translation: The Case of Translating Finnish into English
Gongbo Tang, Rico Sennrich, and Joakim Nivre. 2021 · 2021
Later among the works it cites.
Counterfactual Explanations for Machine Learning: Challenges Revisited
Sahil Verma, John Dickerson, and Keegan Hines. 2021 · 2021
Later among the works it cites.
Polyjuice: Generating counterfactuals for explaining, evaluating, and improving models
Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel S. Weld. 2021a · 2021
Later among the works it cites.
Sakg-bert: Enabling language representation with knowledge graphs for chinese sentiment analysis
Xiaoyan Yan, Fanghong Jian, and Bo Sun. 2021 · 2021
Later among the works it cites.
S-lime: Stabilized-lime for model explanation
Zhengze Zhou, Giles Hooker, and Fei Wang. 2021 · 2021
Later among the works it cites.