Fetching the paper…
Reading the bibliography…
In this paper, we tackle interactive debugging of "gray-box" concept-based models (CBMs).
Interpretability beyond classification output: Semantic Bottleneck Networks
Losch, M.; Fritz, M.; and Schiele, B. 2019 · 1907
Earlier work this paper cites.
Probability product kernels
Jebara, T.; Kondor, R.; and Howard, A. 2004 · 2004
Earlier work this paper cites.
Towards scalable representations of object categories: Learning a hierarchy of parts
Fidler, S.; and Leonardis, A. 2007 · 2007
Earlier work this paper cites.
Machine guides, human supervises: Interactive learning with global explanations
Popordanoska, T.; Kumar, M.; and Teso, S. 2020 · 2009
Earlier work this paper cites.
How to explain individual classification decisions
Baehrens, D.; Schroeter, T.; Harmeling, S.; Kawanabe, M.; Hansen, K.; and Müller, K.-R. 2010 · 2010
Earlier work this paper cites.
This Looks Like That, Because… Explaining Prototypes for Interpretable Image Recognition
Nauta, M.; Jutte, A.; Provoost, J.; and Seifert, C. 2020 · 2011
Earlier work this paper cites.
ProtoPShare: Prototype Sharing for Interpretable Image Classification and Similarity Discovery
Rymarczyk, D.; Struski, Ł.; Tabor, J.; and Zieliński, B. 2020 · 2011
Earlier work this paper cites.
Learning Interpretable Concept-Based Models with Human Feedback
Lage, I.; and Doshi-Velez, F. 2020 · 2012
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C.; Zaremba, W.; Sutskever, I.; Bruna, J.; Erhan, D.; Goodfellow, I.; and Fergus, R. 2013 · 2013
Earlier work this paper cites.
“Why should I trust you?” Explaining the predictions of any classifier
Ribeiro, M. T.; Singh, S.; and Guestrin, C. 2016 · 2016
Earlier work this paper cites.
Right for the right reasons: training differentiable models by constraining their explanations
Ross, A. S.; Hughes, M. C.; and Doshi-Velez, F. 2017 · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Sundararajan, M.; Taly, A.; and Yan, Q. 2017 · 2017
Earlier work this paper cites.
Towards robust interpretability with self-explaining neural networks
Alvarez-Melis, D.; and Jaakkola, T. S. 2018 · 2018
Earlier work this paper cites.
Deep learning for case-based reasoning through prototypes: A neural network that explains its predictions
Li, O.; Liu, H.; Chen, C.; and Rudin, C. 2018 · 2018
Earlier work this paper cites.
The Mythos of Model Interpretability
Lipton, Z. C. 2018 · 2018
Cited alongside, same era.
This Looks Like That: Deep Learning for Interpretable Image Recognition
Chen, C.; Li, O.; Tao, D.; Barnett, A.; Rudin, C.; and Su, J. K. 2019 · 2019
Cited alongside, same era.
Explanations can be manipulated and geometry is to blame
Dombrowski, A.-K.; Alber, M.; Anders, C.; Ackermann, M.; Müller, K.-R.; and Kessel, P. 2019 · 2019
Cited alongside, same era.
Interpretable image recognition with hierarchical prototypes
Hase, P.; Chen, C.; Li, O.; and Rudin, C. 2019 · 2019
Cited alongside, same era.
Unmasking Clever Hans predictors and assessing what machines really learn
Lapuschkin, S.; Wäldchen, S.; Binder, A.; Montavon, G.; Samek, W.; and Müller, K.-R. 2019 · 2019
Cited alongside, same era.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
FIND: human-in-the-loop debugging deep text classifiers
Lertvittayakumjorn, P.; Specia, L.; and Toni, F. 2020 · 2020
Later among the works it cites.
Making deep neural networks right for the right scientific reasons by interacting with their explanations
Schramowski, P.; Stammer, W.; Teso, S.; Brugger, A.; Herbert, F.; Shao, X.; Luigs, H.-G.; Mahlein, A.-K.; and Kersting, K. 2020 · 2020
Later among the works it cites.
When explanations lie: Why many modified bp attributions fail
Sixt, L.; Granz, M.; and Landgraf, T. 2020 · 2020
Later among the works it cites.
On Completeness-aware Concept-Based Explanations in Deep Neural Networks
Yeh, C.-K.; Kim, B.; Arik, S.; Li, C.-L.; Pfister, T.; and Ravikumar, P. 2020 · 2020
Later among the works it cites.
Debiasing Concept-based Explanations with Causal Analysis
Bahadori, M. T.; and Heckerman, D. 2021 · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rudin, C. 2019 · 2019
Cited alongside, same era.
Taking a hint: Leveraging explanations to make vision and language models more grounded
Selvaraju, R. R.; Lee, S.; Shen, Y.; Jin, H.; Ghosh, S.; Heck, L.; Batra, D.; and Parikh, D. 2019 · 2019
Cited alongside, same era.
Toward Faithful Explanatory Active Learning with Self-explainable Neural Nets
Teso, S. 2019 · 2019
Cited alongside, same era.
Explanatory interactive machine learning
Teso, S.; and Kersting, K. 2019 · 2019
Cited alongside, same era.
Dissonance between human and machine understanding
Zhang, Z.; Singh, J.; Gadiraju, U.; and Anand, A. 2019 · 2019
Cited alongside, same era.
Contextual Explanation Networks
Al-Shedivat, M.; Dubey, A.; and Xing, E. P. 2020 · 2020
Cited alongside, same era.
Concept whitening for interpretable image recognition
Chen, Z.; Bei, Y.; and Rudin, C. 2020 · 2020
Cited alongside, same era.
Barnett, A. J.; Schwartz, F. R.; Tao, C.; Chen, C.; Ren, Y.; Lo, J. Y.; and Rudin, C. 2021 · 2021
Closest in time.
Knowledge distillation: A survey
Gou, J.; Yu, B.; Maybank, S. J.; and Tao, D. 2021 · 2021
Closest in time.
Hoffmann, A.; Fanconi, C.; Rade, R.; and Kohler, J. 2021 · 2021
Closest in time.
Neural Prototype Trees for Interpretable Fine-grained Image Recognition
Nauta, M.; van Bree, R.; and Seifert, C. 2021 · 2021
Closest in time.
Right for Better Reasons: Training Differentiable Models by Constraining their Influence Function
Shao, X.; Skryagin, A.; Schramowski, P.; Stammer, W.; and Kersting, K. 2021 · 2021
Closest in time.
Right for the Right Concept: Revising Neuro-Symbolic Concepts by Interacting with their Explanations
Stammer, W.; Schramowski, P.; and Kersting, K. 2021 · 2021
Closest in time.
Interactive Label Cleaning with Example-based Explanations
Teso, S.; Bontempelli, A.; Giunchiglia, F.; and Passerini, A. 2021 · 2021
Closest in time.
On the tractability of SHAP explanations
Van den Broeck, G.; Lykov, A.; Schleich, M.; and Suciu, D. 2021 · 2021
Closest in time.
HILDIF: Interactive Debugging of NLI Models Using Influence Functions
Zylberajch, H.; Lertvittayakumjorn, P.; and Toni, F. 2021 · 2021
Closest in time.