Fetching the paper…
Reading the bibliography…
With the increased deployment of machine learning models in various real-world applications, researchers and practitioners alike have emphasized the need for explanations of model behaviour.
Visualizing and understanding convolutional networks
Matthew D Zeiler and Rob Fergus · 2014
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Song Han, Jeff Pool, John Tran, and William Dally · 2015
Earlier work this paper cites.
" why should i trust you?" explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
Evaluating the visualization of what a deep neural network has learned
Wojciech Samek, Alexander Binder, Grégoire Montavon, Sebastian Lapuschkin, and Klaus-Robert Müller · 2016
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee · 2017
Earlier work this paper cites.
Smoothgrad: removing noise by adding noise
Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda Viégas, and Martin Wattenberg · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra · 2017
Earlier work this paper cites.
Interpretable explanations of black boxes by meaningful perturbation
Ruth C Fong and Andrea Vedaldi · 2017
Earlier work this paper cites.
Real time image saliency for black box classifiers
Piotr Dabkowski and Yarin Gal · 2017
Earlier work this paper cites.
Generalized additive models
Trevor J Hastie · 2017
Earlier work this paper cites.
Non-convex optimization for machine learning
Prateek Jain and Purushottam Kar · 2017
Cited alongside, same era.
General data protection regulation (gdpr)
The European Commission · 2018
Cited alongside, same era.
Learning to explain: An information-theoretic perspective on model interpretation
Jianbo Chen, Le Song, Martin Wainwright, and Michael Jordan · 2018
Cited alongside, same era.
Labeled optical coherence tomography (oct) and chest x-ray images for classification
Daniel Kermany, Kang Zhang, Michael Goldbaum, et al · 2018
Cited alongside, same era.
Large-scale celebfaces attributes (celeba) dataset
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang · 2018
Cited alongside, same era.
A benchmark for interpretability methods in deep neural networks
Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim · 2019
Cited alongside, same era.
Fooling neural network interpretations via adversarial model manipulation
Juyeon Heo, Sunghwan Joo, and Taesup Moon · 2019
Later among the works it cites.
Concept bottleneck models
Pang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann, Emma Pierson, Been Kim, and Percy Liang · 2020
Later among the works it cites.
" how do i fool you?" manipulating user trust via misleading black box explanations
Himabindu Lakkaraju and Osbert Bastani · 2020
Later among the works it cites.
Do input gradients highlight discriminative features?
Harshay Shah, Prateek Jain, and Praneeth Netrapalli · 2021
Later among the works it cites.
Have we learned to explain?: How interpretability methods can learn to encode predictions in their interpretations
Neil Jethani, Mukund Sudarshan, Yindalon Aphinyanaphongs, and Rajesh Ranganath · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Full-gradient representation for neural network visualization
Suraj Srinivas and François Fleuret · 2019
Cited alongside, same era.
Understanding deep networks via extremal perturbations and smooth masks
Ruth Fong, Mandela Patrick, and Andrea Vedaldi · 2019
Cited alongside, same era.
Invase: Instance-wise variable selection using neural networks
Jinsung Yoon, James Jordon, and Mihaela van der Schaar · 2019
Cited alongside, same era.
This looks like that: deep learning for interpretable image recognition
Chaofan Chen, Oscar Li, Daniel Tao, Alina Barnett, Cynthia Rudin, and Jonathan K Su · 2019
Cited alongside, same era.
The White House · 2022
Later among the works it cites.
Tessa Han, Suraj Srinivas, and Himabindu Lakkaraju · 2022
Later among the works it cites.
B-cos networks: Alignment is all we need for interpretability
Moritz Böhle, Mario Fritz, and Bernt Schiele · 2022
Later among the works it cites.
Openxai: Towards a transparent evaluation of model explanations
Chirag Agarwal, Satyapriya Krishna, Eshika Saxena, Martin Pawelczyk, Nari Johnson, Isha Puri, Marinka Zitnik, and Himabindu Lakkaraju · 2022
Later among the works it cites.