Fetching the paper…
Reading the bibliography…
Feature attribution methods are popular for explaining neural network predictions, and they are often evaluated on metrics such as comprehensiveness and sufficiency.
One explanation does not fit all: A toolkit and taxonomy of AI explainability techniques
Vijay Arya, Rachel KE Bellamy, Pin-Yu Chen, Amit Dhurandhar, Michael Hind, Samuel C Hoffman, Stephanie Houde, Q Vera Liao, Ronny Luss, Aleksandra Mojsilović, et al. 2019 · 1909
Earlier work this paper cites.
Towards a unified evaluation of explanation methods without ground truth
Hao Zhang, Jiayi Chen, Haotian Xue, and Quanshi Zhang. 2019 · 1911
Earlier work this paper cites.
Optimization by simulated annealing
Scott Kirkpatrick, C Daniel Gelatt Jr, and Mario P Vecchi. 1983 · 1983
Earlier work this paper cites.
The Shapley Value: Essays in Honor of Lloyd S. Shapley
Alvin E Roth. 1988 · 1988
Earlier work this paper cites.
UCI Adult data set
Ronny Kohavi and Barry Becker. 1996 · 1996
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2013 · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D Zeiler and Rob Fergus. 2014 · 2014
Earlier work this paper cites.
Yelp dataset challenge: Review rating prediction
Nabiha Asghar. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
"Why should I trust you?" explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Evaluating the visualization of what a deep neural network has learned
Wojciech Samek, Alexander Binder, Grégoire Montavon, Sebastian Lapuschkin, and Klaus-Robert Müller. 2016 · 2016
Earlier work this paper cites.
Real time image saliency for black box classifiers
Piotr Dabkowski and Yarin Gal. 2017 · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee. 2017 · 2017
Cited alongside, same era.
SmoothGrad: Removing noise by adding noise
Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda Viégas, and Martin Wattenberg. 2017 · 2017
Cited alongside, same era.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Cited alongside, same era.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim. 2018 · 2018
Cited alongside, same era.
Pathologies of neural models make interpretations difficult
Shi Feng, Eric Wallace, Alvin Grissom II, Mohit Iyyer, Pedro Rodriguez, and Jordan Boyd-Graber. 2018 · 2018
Cited alongside, same era.
Towards deep learning models resistant to adversarial attacks
Shortcut learning in deep neural networks
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann. 2020 · 2020
Later among the works it cites.
Interpretation of NLP models through input marginalization
Siwon Kim, Jihun Yi, Eunji Kim, and Sungroh Yoon. 2020 · 2020
Later among the works it cites.
Fooling LIME and SHAP: Adversarial attacks on post hoc explanation methods
Dylan Slack, Sophie Hilgard, Emily Jia, Sameer Singh, and Himabindu Lakkaraju. 2020 · 2020
Later among the works it cites.
Does the whole exceed its parts? the effect of AI explanations on complementary team performance
Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld. 2021 · 2021
Later among the works it cites.
Improving the faithfulness of attention-based explanations with task-specific information for text classification
George Chrysostomou and Nikolaos Aletras. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2018 · 2018
Cited alongside, same era.
A theoretical explanation for perplexing behaviors of backpropagation-based visualizations
Weili Nie, Yang Zhang, and Ankit Patel. 2018 · 2018
Cited alongside, same era.
RISE: Randomized input sampling for explanation of black-box models
Vitali Petsiuk, Abir Das, and Kate Saenko. 2018 · 2018
Cited alongside, same era.
Model agnostic supervised local explanations
Gregory Plumb, Denali Molitor, and Ameet S Talwalkar. 2018 · 2018
Cited alongside, same era.
Interpretation of neural networks is fragile
Amirata Ghorbani, Abubakar Abid, and James Zou. 2019 · 2019
Cited alongside, same era.
Is attention interpretable?
Sofia Serrano and Noah A. Smith. 2019 · 2019
Cited alongside, same era.
AllenNLP Interpret: A framework for explaining predictions of NLP models
Eric Wallace, Jens Tuyls, Junlin Wang, Sanjay Subramanian, Matt Gardner, and Sameer Singh. 2019 · 2019
Cited alongside, same era.
Leveraging latent features for local explanations
Ronny Luss, Pin-Yu Chen, Amit Dhurandhar, Prasanna Sattigeri, Yunfeng Zhang, Karthikeyan Shanmugam, and Chun-Chen Tu. 2021 · 2021
Later among the works it cites.
Towards understanding the behaviors of optimal deep active learning algorithms
Yilun Zhou, Adithya Renduchintala, Xian Li, Sida Wang, Yashar Mehdad, and Asish Ghoshal. 2021 · 2021
Later among the works it cites.
A protocol for evaluating the faithfulness of input salience methods for text classification
Jasmijn Bastings, Sebastian Ebert, Polina Zablotskaia, Anders Sandholm, and Katja Filippova. 2022 · 2022
Closest in time.
A comparative study of faithfulness metrics for model interpretability methods
Chun Sik Chan, Huanqi Kong, and Liang Guanqing. 2022 · 2022
Closest in time.
Interpretable machine learning: Moving from mythos to diagnostics
Valerie Chen, Jeffrey Li, Joon Sik Kim, Gregory Plumb, and Ameet Talwalkar. 2022 · 2022
Closest in time.
The role of explainability in assuring safety of machine learning in healthcare
Yan Jia, John McDermid, Tom Lawton, and Ibrahim Habli. 2022 · 2022
Closest in time.
Logic traps in evaluating attribution scores
Yiming Ju, Yuanzhe Zhang, Zhao Yang, Zhongtao Jiang, Kang Liu, and Jun Zhao. 2022 · 2022
Closest in time.
Double trouble: How to not explain a text classifier’s decisions using counterfactuals synthesized by masked language models?
Thang Pham, Trung Bui, Long Mai, and Anh Nguyen. 2022 · 2022
Closest in time.
The irrationality of neural rationale models
Yiming Zheng, Serena Booth, Julie Shah, and Yilun Zhou. 2022 · 2022
Closest in time.