Fetching the paper…
Reading the bibliography…
Interpretability methods are developed to understand the working mechanisms of black-box models, which is crucial to their responsible deployment.
RoBERTa: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Interpretable decision sets: A joint framework for description and prediction
Himabindu Lakkaraju, Stephen H Bach, and Jure Leskovec. 2016 · 2016
Earlier work this paper cites.
Rationalizing neural predictions
Tao Lei, Regina Barzilay, and Tommi Jaakkola. 2016 · 2016
Earlier work this paper cites.
Visualizing and understanding neural models in NLP
Jiwei Li, Xinlei Chen, Eduard Hovy, and Dan Jurafsky. 2016 · 2016
Earlier work this paper cites.
“Why should I trust you?” Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Evaluating the visualization of what a deep neural network has learned
Wojciech Samek, Alexander Binder, Grégoire Montavon, Sebastian Lapuschkin, and Klaus-Robert Müller. 2016 · 2016
Earlier work this paper cites.
Interpretable explanations of black boxes by meaningful perturbation
Ruth C Fong and Andrea Vedaldi. 2017 · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee. 2017 · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim. 2018 · 2018
Earlier work this paper cites.
Flask web development: Developing web applications with Python
Miguel Grinberg. 2018 · 2018
Earlier work this paper cites.
A theoretical explanation for perplexing behaviors of backpropagation-based visualizations
Weili Nie, Yang Zhang, and Ankit Patel. 2018 · 2018
Cited alongside, same era.
Anchors: High-precision model-agnostic explanations
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2018 · 2018
Cited alongside, same era.
Evaluating recurrent neural network explanations
Leila Arras, Ahmed Osman, Klaus-Robert Müller, and Wojciech Samek. 2019 · 2019
Cited alongside, same era.
Interpretable neural predictions with differentiable binary variables
Jasmijn Bastings, Wilker Aziz, and Ivan Titov. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Interpretation of neural networks is fragile
Amirata Ghorbani, Abubakar Abid, and James Zou. 2019 · 2019
Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?
Alon Jacovi and Yoav Goldberg. 2020 · 2020
Later among the works it cites.
Does the whole exceed its parts? The effect of AI explanations on complementary team performance
Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld. 2021 · 2021
Later among the works it cites.
StereoSet: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy. 2021 · 2021
Later among the works it cites.
Explaining NLP models via minimal contrastive editing (MiCE)
Alexis Ross, Ana Marasović, and Matthew Peters. 2021 · 2021
Later among the works it cites.
Low frequency names exhibit bias and overfitting in contextualizing language models
Robert Wolfe and Aylin Caliskan. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2019 · 2019
Cited alongside, same era.
Faithful and customizable explanations of black box models
Himabindu Lakkaraju, Ece Kamar, Rich Caruana, and Jure Leskovec. 2019 · 2019
Cited alongside, same era.
Understanding global feature contributions with additive importance measures
Ian Covert, Scott M Lundberg, and Su-In Lee. 2020 · 2020
Cited alongside, same era.
Evaluating explainable AI: Which algorithmic explanations help users predict model behavior?
Peter Hase and Mohit Bansal. 2020 · 2020
Cited alongside, same era.
On the robustness of interpretability methods
David Alvarez-Melis and Tommi S Jaakkola. 2018a
Cited in the paper.
Towards robust interpretability with self-explaining neural networks
David Alvarez-Melis and Tommi S Jaakkola. 2018b
Cited in the paper.
Polyjuice: Generating counterfactuals for explaining, evaluating, and improving models
Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel Weld. 2021 · 2021
Later among the works it cites.
The irrationality of neural rationale models
Yiming Zheng, Serena Booth, Julie Shah, and Yilun Zhou. 2021 · 2021
Later among the works it cites.
Post hoc explanations may be ineffective for detecting unknown spurious correlation
Julius Adebayo, Michael Muelly, Harold Abelson, and Been Kim. 2022 · 2022
Closest in time.
Evaluating Explanations: How Much Do Explanations from the Teacher Aid Students?
Danish Pruthi, Rachit Bansal, Bhuwan Dhingra, Livio Baldini Soares, Michael Collins, Zachary C. Lipton, Graham Neubig, and William W. Cohen. 2022 · 2022
Closest in time.
Do feature attribution methods correctly attribute features?
Yilun Zhou, Serena Booth, Marco Tulio Ribeiro, and Julie Shah. 2022 · 2022
Closest in time.