Fetching the paper…
Reading the bibliography…
With the growing popularity of deep-learning based NLP models, comes a need for interpretable systems.
Interpretable machine learning: definitions, methods, and applications
W. James Murdoch, Chandan Singh, Karl Kumbier, Reza Abbasi-Asl, and Bin Yu. 2019 · 1901
Earlier work this paper cites.
An evaluation of the human-interpretability of explanation
Isaac Lage, Emily Chen, Jeffrey He, Menaka Narayanan, Been Kim, Sam Gershman, and Finale Doshi-Velez. 2019 · 1902
Earlier work this paper cites.
Do human rationales improve machine explanations?
Julia Strout, Ye Zhang, and Raymond J. Mooney. 2019 · 1905
Earlier work this paper cites.
Explain yourself! leveraging language models for commonsense reasoning
Nazneen Fatema Rajani, Bryan McCann, Caiming Xiong, and Richard Socher. 2019 · 1906
Earlier work this paper cites.
Analyzing the structure of attention in a transformer language model
Jesse Vig and Yonatan Belinkov. 2019 · 1906
Earlier work this paper cites.
Do transformer attention heads provide transparency in abstractive summarization?
Joris Baan, Maartje ter Hoeve, Marlies van der Wees, Anne Schuth, and Maarten de Rijke. 2019 · 1907
Earlier work this paper cites.
A human-grounded evaluation of SHAP for alert processing
Hilde J. P. Weerts, Werner van Ipenburg, and Mykola Pechenizkiy. 2019 · 1907
Earlier work this paper cites.
Learning to deceive with attention-based explanations
Danish Pruthi, Mansi Gupta, Bhuwan Dhingra, Graham Neubig, and Zachary C. Lipton. 2019 · 1909
Earlier work this paper cites.
Attention interpretability across NLP tasks
Shikhar Vashishth, Shyam Upadhyay, Gaurav Singh Tomar, and Manaal Faruqui. 2019 · 1909
Earlier work this paper cites.
Rethinking cooperative rationalization: Introspective extraction and complement control
Mo Yu, Shiyu Chang, Yang Zhang, and Tommi S. Jaakkola. 2019 · 1910
Earlier work this paper cites.
Harvey Friedman’s Research on the Foundations of Mathematics
L.A. Harrington, M.D. Morley, A. Šcedrov, and S.G. Simpson. 1985 · 1985
Earlier work this paper cites.
Modeling annotators: A generative approach to learning from annotator rationales
Omar Zaidan and Jason Eisner. 2008 · 2008
Earlier work this paper cites.
Towards extracting faithful and descriptive representations of latent variable models
Vicente Iván Sánchez Carmona, Tim Rocktäschel, Sebastian Riedel, and Sameer Singh. 2015 · 2015
Earlier work this paper cites.
”what is relevant in a text document?”: An interpretable machine learning approach
Leila Arras, Franziska Horn, Grégoire Montavon, Klaus-Robert Müller, and Wojciech Samek. 2016 · 2016
Earlier work this paper cites.
Yes, we care! results of the ethics and natural language processing surveys
Karën Fort and Alain Couillault. 2016 · 2016
Earlier work this paper cites.
“Why Should I Trust You?”: Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Interpretability of deep learning models: a survey of results
Supriyo Chakraborty, Richard Tomsett, Ramya Raghavendra, Daniel Harborne, Moustafa Alzantot, Federico Cerutti, Mani Srivastava, Alun Preece, Simon Julier, Raghuveer M Rao, et al. 2017 · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim. 2017 · 2017
Cited alongside, same era.
The promise and peril of human evaluation for model interpretability
Bernease Herman. 2017 · 2017
Cited alongside, same era.
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, and Rory Sayres. 2017 · 2017
Cited alongside, same era.
Interactive visualization and manipulation of attention-based neural machine translation
Jaesong Lee, Joong-Hwi Shin, and Jun-Seok Kim. 2017 · 2017
Cited alongside, same era.
A unified approach to interpreting model predictions
Scott M. Lundberg and Su-In Lee. 2017 · 2017
Evaluating neural network explanation methods using hybrid documents and morphological prediction
Nina Pörner, Hinrich Schütze, and Benjamin Roth. 2018 · 2018
Later among the works it cites.
Please stop explaining black box models for high stakes decisions
Cynthia Rudin. 2018 · 2018
Later among the works it cites.
Rule induction for global explanation of trained models
Madhumita Sushil, Simon Šuster, and Walter Daelemans. 2018 · 2018
Later among the works it cites.
Faithful multimodal explanation for visual question answering
Jialin Wu and Raymond J. Mooney. 2018 · 2018
Later among the works it cites.
Looking deeper into deep learning model: Attribution-based explanations of textcnn
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Trends and trajectories for explainable, accountable and intelligible systems: An HCI research agenda
Ashraf M. Abdul, Jo Vermeulen, Danding Wang, Brian Y. Lim, and Mohan S. Kankanhalli. 2018 · 2018
Cited alongside, same era.
On the robustness of interpretability methods
David Alvarez-Melis and Tommi S. Jaakkola. 2018 · 2018
Cited alongside, same era.
DALEX: explainers for complex predictive models in R
Przemyslaw Biecek. 2018 · 2018
Cited alongside, same era.
Pathologies of neural models make interpretation difficult
Shi Feng, Eric Wallace, Alvin Grissom II, Mohit Iyyer, Pedro Rodriguez, and Jordan L. Boyd-Graber. 2018 · 2018
Cited alongside, same era.
Interpreting recurrent and attention-based neural models: a case study on natural language inference
Reza Ghaeini, Xiaoli Z. Fern, and Prasad Tadepalli. 2018 · 2018
Cited alongside, same era.
Explaining explanations: An overview of interpretability of machine learning
Leilani H Gilpin, David Bau, Ben Z Yuan, Ayesha Bajwa, Michael Specter, and Lalana Kagal. 2018 · 2018
Cited alongside, same era.
A survey of methods for explaining black box models
Riccardo Guidotti, Anna Monreale, Salvatore Ruggieri, Franco Turini, Fosca Giannotti, and Dino Pedreschi. 2018 · 2018
Cited alongside, same era.
Wenting Xiong, Iftitahu Ni’mah, Juan M. G. Huesca, Werner van Ipenburg, Jan Veldsink, and Mykola Pechenizkiy. 2018 · 2018
Later among the works it cites.
Can i trust the explainer? verifying post-hoc explanatory methods
Oana-Maria Camburu, Eleonora Giunchiglia, Jakob Foerster, Thomas Lukasiewicz, and Phil Blunsom. 2019 · 2019
Later among the works it cites.
Eraser: A benchmark to evaluate rationalized nlp models
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace. 2019 · 2019
Later among the works it cites.
What can ai do for me? evaluating machine learning interpretations in cooperative play
Shi Feng and Jordan Boyd-Graber. 2019 · 2019
Later among the works it cites.
Interpretation of neural networks is fragile
Amirata Ghorbani, Abubakar Abid, and James Zou. 2019 · 2019
Later among the works it cites.
Attention is not explanation
Sarthak Jain and Byron C. Wallace. 2019 · 2019
Later among the works it cites.
Interpreting interpretability: Understanding data scientists’ use of interpretability tools for machine learning
Harmanpreet Kaur, Harsha Nori, Samuel Jenkins, Rich Caruana, Hanna M. Wallach, and Jennifer Wortman Vaughan. 2019 · 2019
Later among the works it cites.
The (un)reliability of saliency methods
Pieter-Jan Kindermans, Sara Hooker, Julius Adebayo, Maximilian Alber, Kristof T. Schütt, Sven Dähne, Dumitru Erhan, and Been Kim. 2019 · 2019
Later among the works it cites.
Faithful and customizable explanations of black box models
Himabindu Lakkaraju, Ece Kamar, Rich Caruana, and Jure Leskovec. 2019 · 2019
Later among the works it cites.
Is attention interpretable?
Sofia Serrano and Noah A. Smith. 2019 · 2019
Later among the works it cites.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter. 2019 · 2019
Later among the works it cites.
A formal approach to explainability
Lior Wolf, Tomer Galanti, and Tamir Hazan. 2019 · 2019
Later among the works it cites.