Fetching the paper…
Reading the bibliography…
Concept-based interpretability methods aim to explain deep neural network model predictions using a predefined set of semantic concepts.
The Pascal Visual Object Classes (VOC) Challenge
M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman · 2010
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Earlier work this paper cites.
The Caltech-UCSD Birds-200-2011 Dataset
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie · 2011
Earlier work this paper cites.
Diagnosing error in object detectors
Derek Hoiem, Yodsawalai Chodpathumwan, and Qieyun Dai · 2012
Earlier work this paper cites.
Intrinsic images in the wild
Sean Bell, Kavita Bala, and Noah Snavely · 2014
Earlier work this paper cites.
Describing textures in the wild
Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi · 2014
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D Zeiler and Rob Fergus · 2014
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Salient deconvolutional networks
Aravindh Mahendran and Andrea Vedaldi · 2016
Earlier work this paper cites.
Learning deep features for discriminative localization
Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba · 2016
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations
David Bau, Bolei Zhou, Aditya Khosla, Aude Oliva, and Antonio Torralba · 2017
Earlier work this paper cites.
Grad-CAM: Visual explanations from deep networks via gradient-based localization
Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra · 2017
Earlier work this paper cites.
Places: A 10 million image database for scene recognition
Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba · 2017
Earlier work this paper cites.
Scene parsing through ADE20k dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim · 2018
Earlier work this paper cites.
Approximating CNNs with bag-of-local-features models works surprisingly well on imagenet
Wieland Brendel and Matthias Bethge · 2018
Earlier work this paper cites.
Grad-CAM++: Generalized gradient-based visual explanations for deep convolutional networks
Aditya Chattopadhay, Anirban Sarkar, Prantik Howlader, and Vineeth N Balasubramanian · 2018
Earlier work this paper cites.
Net2Vec: Quantifying and explaining how concepts are encoded by filters in deep neural networks
Ruth Fong and Andrea Vedaldi · 2018
Earlier work this paper cites.
Explaining explanations: An overview of interpretability of machine learning
Leilani H. Gilpin, David Bau, Ben Z Yuan, Ayesha Bajwa, Michael Specter, and Lalana Kagal · 2018
Cited alongside, same era.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al · 2018
Cited alongside, same era.
RISE: Randomized input sampling for explanation of black-box models
Vitali Petsiuk, Abir Das, and Kate Saenko · 2018
Cited alongside, same era.
Top-down neural attention by excitation backprop
Jianming Zhang, Sarah Adel Bargal, Zhe Lin, Jonathan Brandt, Xiaohui Shen, and Stan Sclaroff · 2018
Cited alongside, same era.
Visual interpretability for deep learning: A survey
Quanshi Zhang and Song-Chun Zhu · 2018
Cited alongside, same era.
Interpretable basis decomposition for visual explanation
How can i explain this to you? an empirical study of deep neural network explanation methods
Jeya Vikranth Jeyakumar, Joseph Noor, Yu-Hsi Cheng, Luis Garcia, and Mani Srivastava · 2020
Later among the works it cites.
Concept bottleneck models
Pang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann, Emma Pierson, Been Kim, and Percy Liang · 2020
Later among the works it cites.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Later among the works it cites.
Why do these match? explaining the behavior of image similarity models
Bryan A. Plummer, Mariya I. Vasileva, Vitali Petsiuk, Kate Saenko, and David Forsyth · 2020
Later among the works it cites.
There and back again: Revisiting backpropagation saliency methods
Sylvestre-Alvise Rebuffi, Ruth Fong, Xu Ji, and Andrea Vedaldi · 2020
Later among the works it cites.
Explainability fact sheets: A framework for systematic assessment of explainable approaches
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bolei Zhou, Yiyou Sun, David Bau, and Antonio Torralba · 2018
Cited alongside, same era.
This looks like that: Deep learning for interpretable image recognition
Chaofan Chen, Oscar Li, Chaofan Tao, Alina Jade Barnett, Jonathan Su, and Cynthia Rudin · 2019
Cited alongside, same era.
Understanding deep networks via extremal perturbations and smooth masks
Ruth Fong, Mandela Patrick, and Andrea Vedaldi · 2019
Cited alongside, same era.
Counterfactual visual explanations
Yash Goyal, Ziyan Wu, Jan Ernst, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Cited alongside, same era.
A benchmark for interpretability methods in deep neural networks
Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim · 2019
Cited alongside, same era.
The (un) reliability of saliency methods
Pieter-Jan Kindermans, Sara Hooker, Julius Adebayo, Maximilian Alber, Kristof T Schütt, Sven Dähne, Dumitru Erhan, and Been Kim · 2019
Cited alongside, same era.
Human evaluation of models built for interpretability
Isaac Lage, Emily Chen, Jeffrey He, Menaka Narayanan, Been Kim, Samuel J Gershman, and Finale Doshi-Velez · 2019
Cited alongside, same era.
Kacper Sokol and Peter Flach · 2020
Later among the works it cites.
On completeness-aware concept-based explanations in deep neural networks
Chih-Kuan Yeh, Been Kim, Sercan Arik, Chun-Liang Li, Tomas Pfister, and Pradeep Ravikumar · 2020
Later among the works it cites.
An interpretability illusion for bert, 2021
Tolga Bolukbasi, Adam Pearce, Ann Yuan, Andy Coenen, Emily Reif, Fernanda Viégas, and Martin Wattenberg · 2021
Later among the works it cites.
Datasheets for datasets
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé Iii, and Kate Crawford · 2021
Later among the works it cites.
This looks like that… does it? Shortcomings of latent space prototype interpretability in deep networks
Adrian Hoffmann, Claudio Fanconi, Rahul Rade, and Jonas Kohler · 2021
Later among the works it cites.
Do concept bottleneck models learn as intended?
Andrei Margeloiu, Matthew Ashman, Umang Bhatt, Yanzhi Chen, Mateja Jamnik, and Adrian Weller · 2021
Later among the works it cites.
Neural prototype trees for interpretable fine-grained image recognition
Meike Nauta, Ron van Bree, and Christin Seifert · 2021
Later among the works it cites.
Fair attribute classification through latent space de-biasing
Vikram V. Ramaswamy, Sunnie S. Y. Kim, and Olga Russakovsky · 2021
Later among the works it cites.
Interpretable machine learning: Fundamental principles and 10 grand challenges
Cynthia Rudin, Chaofan Chen, Zhi Chen, Haiyang Huang, Lesia Semenova, and Chudi Zhong · 2021
Later among the works it cites.
Post hoc explanations may be ineffective for detecting unknown spurious correlation
Julius Adebayo, Michael Muelly, Harold Abelson, and Been Kim · 2022
Closest in time.
Interactive model cards: A human-centered approach to model documentation
Anamaria Crisan, Margaret Drouhard, Jesse Vig, and Nazneen Rajani · 2022
Closest in time.
HIVE: Evaluating the human interpretability of visual explanations
Sunnie S. Y. Kim, Nicole Meister, Vikram V. Ramaswamy, Ruth Fong, and Olga Russakovsky · 2022
Closest in time.
Vikram V. Ramaswamy, Sunnie S. Y. Kim, Nicole Meister, Ruth Fong, and Olga Russakovsky · 2022
Closest in time.
“Help me help the AI”: Understanding how explainability can support human-AI interaction
Sunnie S. Y. Kim, Elizabeth Anne Watkins, Olga Russakovsky, Ruth Fong, and Andrés Monroy-Hernández · 2023
Closest in time.