Fetching the paper…
Reading the bibliography…
Human explanations of high-level decisions are often expressed in terms of key concepts the decisions are based on.
Explaining classifiers with causal concept effect (cace)
Yash Goyal, Amir Feder, Uri Shalit, and Been Kim · 1907
Earlier work this paper cites.
What some concepts might not be
Sharon Lee Armstrong, Lila R. Gleitman, and Henry Gleitman · 1983
Earlier work this paper cites.
A value for n-person games , page 31–40
Lloyd S. Shapley · 1988
Earlier work this paper cites.
A Bayesian framework for concept learning
Joshua Brett Tenenbaum · 1999
Earlier work this paper cites.
Latent dirichlet allocation
David M Blei, Andrew Y Ng, and Michael I Jordan · 2003
Earlier work this paper cites.
Axiomatic characterizations of probabilistic and cardinal-probabilistic interaction indices
Katsushige Fujimoto, Ivan Kojadinovic, and Jean-Luc Marichal · 2006
Earlier work this paper cites.
Supervised topic models
Jon D Mcauliffe and David M Blei · 2008
Earlier work this paper cites.
Learning to detect unseen object classes by between-class attribute transfer
Christoph H Lampert, Hannes Nickisch, and Stefan Harmeling · 2009
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Pcanet: A simple deep learning baseline for image classification?
Tsung-Han Chan, Kui Jia, Shenghua Gao, Jiwen Lu, Zinan Zeng, and Yi Ma · 2015
Earlier work this paper cites.
Why should i trust you?: Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
Evaluating the visualization of what a deep neural network has learned
Wojciech Samek, Alexander Binder, Grégoire Montavon, Sebastian Lapuschkin, and Klaus-Robert Müller · 2016
Earlier work this paper cites.
Examples are not enough, learn to criticize! criticism for interpretability
Been Kim, Rajiv Khanna, and Oluwasanmi O Koyejo · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee · 2017
Earlier work this paper cites.
Smoothgrad: removing noise by adding noise
Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda Viégas, and Martin Wattenberg · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang · 2017
Cited alongside, same era.
Counterfactual explanations without opening the black box: Automated decisions and the gdpr
S. Wachter, Brent D. Mittelstadt, and Chris Russell · 2017
Cited alongside, same era.
Towards better understanding of gradient-based attribution methods for deep neural networks
Marco Ancona, Enea Ceolini, Cengiz Öztireli, and Markus Gross · 2017
Cited alongside, same era.
Learning to generate reviews and discovering sentiment
Alec Radford, Rafal Józefowicz, and Ilya Sutskever · 2017
Cited alongside, same era.
Simple black-box adversarial attacks on deep neural networks
N. Narodytska and S. Kasiviswanathan · 2017
Cited alongside, same era.
Towards automatic concept-based explanations
Amirata Ghorbani, James Wexler, James Zou, and Been Kim · 2019
Closest in time.
Educe: Explaining model decisions through unsupervised concepts extraction
Diane Bouchacourt and Ludovic Denoyer · 2019
Closest in time.
Evaluating explanation without ground truth in interpretable machine learning
Fan Yang, Mengnan Du, and Xia Hu · 2019
Closest in time.
Interpreting black box predictions using fisher kernels
Rajiv Khanna, Been Kim, Joydeep Ghosh, and Sanmi Koyejo · 2019
Closest in time.
Sercan Ö. Arik and Tomas Pfister · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al · 2018
Cited alongside, same era.
Interpretable basis decomposition for visual explanation
Bolei Zhou, Yiyou Sun, David Bau, and Antonio Torralba · 2018
Cited alongside, same era.
Explaining explanations: An overview of interpretability of machine learning
Leilani H Gilpin, David Bau, Ben Z Yuan, Ayesha Bajwa, Michael Specter, and Lalana Kagal · 2018
Cited alongside, same era.
L-shapley and c-shapley: Efficient model interpretation for structured data
Jianbo Chen, Le Song, Martin J. Wainwright, and Michael I. Jordan · 2018
Cited alongside, same era.
Representer point selection for explaining deep neural networks
Chih-Kuan Yeh, Joon Kim, Ian En-Hsu Yen, and Pradeep K Ravikumar · 2018
Cited alongside, same era.
Explanations based on the missing: Towards contrastive explanations with pertinent negatives
Amit Dhurandhar, Pin-Yu Chen, Ronny Luss, Chun-Chen Tu, Paishun Ting, Karthikeyan Shanmugam, and Payel Das · 2018
Cited alongside, same era.
Grounding visual explanations
Lisa Anne Hendricks, Ronghang Hu, Trevor Darrell, and Zeynep Akata · 2018
Cited alongside, same era.
S. Joshi, O. Koyejo, Warut D. Vijitbenjaronk, Been Kim, and Joydeep Ghosh · 2019
Closest in time.
Human evaluation of models built for interpretability
Isaac Lage, Emily Chen, Jeffrey He, Menaka Narayanan, Been Kim, Samuel J Gershman, and Finale Doshi-Velez · 2019
Closest in time.
On the (in)fidelity and sensitivity of explanations
Chih-Kuan Yeh, Cheng-Yu Hsieh, Arun Sai Suggala, David I. Inouye, and Pradeep Ravikumar · 2019
Closest in time.
Benchmarking attribution methods with relative feature importance
Mengjiao Yang and Been Kim · 2019
Closest in time.
Unsupervised speech representation learning using wavenet autoencoders
Jan Chorowski, Ron J Weiss, Samy Bengio, and Aäron van den Oord · 2019
Closest in time.
Beef: Balanced english explanations of forecasts
S. Grover, C. Pulice, G. I. Simari, and V. S. Subrahmanian · 2019
Closest in time.
This looks like that: deep learning for interpretable image recognition
Chaofan Chen, Oscar Li, Daniel Tao, Alina Barnett, Cynthia Rudin, and Jonathan K Su · 2019
Closest in time.
Computing receptive fields of convolutional neural networks
André Araujo, Wade Norris, and Jack Sim · 2019
Closest in time.
Face: Feasible and actionable counterfactual explanations
Rafael Poyiadzi, Kacper Sokol, Raúl Santos-Rodríguez, T. D. Bie, and Peter A. Flach · 2020
Closest in time.
Pang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann, Emma Pierson, Been Kim, and Percy Liang · 2020
Closest in time.