Fetching the paper…
Reading the bibliography…
Explanations have gained an increasing level of interest in the AI and Machine Learning (ML) communities in order to improve model transparency and allow users to form a mental model of a trained ML model.
Explanation-based learning: An alternative view
G. DeJong and R. Mooney · 1986
Earlier work this paper cites.
Explanation-based generalization: A unifying view
T. M. Mitchell, R. M. Keller, and S. T. Kedar-Cabelli · 1986
Earlier work this paper cites.
Extracting tree-structured representations of trained networks
M. Craven and J. Shavlik · 1995
Earlier work this paper cites.
No free lunch theorems for optimization
D. H. Wolpert and W. G. Macready · 1997
Earlier work this paper cites.
Interactive machine learning: letting users build classifiers
M. Ware, E. Frank, G. Holmes, M. Hall, and I. H. Witten · 2001
Earlier work this paper cites.
Training invariant support vector machines
D. DeCoste and B. Schölkopf · 2002
Earlier work this paper cites.
The role of transparency in recommender systems
R. R. Sinha and K. Swearingen · 2002
Earlier work this paper cites.
Interactive machine learning
J. A. Fails and D. R. Olsen Jr · 2003
Earlier work this paper cites.
The structure and function of explanations
T. Lombrozo · 2006
Earlier work this paper cites.
Active learning with feedback on features and instances
H. Raghavan, O. Madani, and R. Jones · 2006
Earlier work this paper cites.
An interactive algorithm for asking and incorporating feature feedback into support vector machines
H. Raghavan and J. Allan · 2007
Earlier work this paper cites.
Toward harnessing user feedback for machine learning
S. Stumpf, V. Rajaram, L. Li, M. Burnett, T. Dietterich, E. Sullivan, R. Drummond, and J. Herlocker · 2007
Earlier work this paper cites.
Effective explanations of recommendations: User-centered design
N. Tintarev and J. Masthoff · 2007
Earlier work this paper cites.
Using “annotator rationales” to improve machine learning for text categorization
O. Zaidan, J. Eisner, and C. Piatko · 2007
Earlier work this paper cites.
Learning from labeled features using generalized expectation criteria
G. Druck, G. Mann, and A. McCallum · 2008
Earlier work this paper cites.
Active learning by labeling features
G. Druck, B. Settles, and A. McCallum · 2009
Earlier work this paper cites.
Causality
J. Pearl · 2009
Earlier work this paper cites.
A unified approach to active dual supervision for labeling features and examples
J. Attenberg, P. Melville, and F. Provost · 2010
Earlier work this paper cites.
How to explain individual classification decisions
D. Baehrens, T. Schroeter, S. Harmeling, M. Kawanabe, K. Hansen, and K.-R. Müller · 2010
Earlier work this paper cites.
Explanatory debugging: Supporting end-user debugging of machine-learned programs
T. Kulesza, S. Stumpf, M. Burnett, W.-K. Wong, Y. Riche, T. Moore, I. Oberst, A. Shinsel, and K. McIntosh · 2010
Earlier work this paper cites.
Prototype selection for interpretable classification
J. Bien and R. Tibshirani · 2011
Earlier work this paper cites.
Closing the loop: Fast, interactive semi-supervised annotation with queries on features and instances
B. Settles · 2011
Earlier work this paper cites.
The constrained weight space svm: learning with ranked features
K. Small, B. C. Wallace, C. E. Brodley, and T. A. Trikalinos · 2011
Earlier work this paper cites.
Critiquing-based recommenders: survey and emerging trends
L. Chen and P. Pu · 2012
Earlier work this paper cites.
Neural-symbolic learning systems: foundations and applications
A. S. d. Garcez, K. B. Broda, and D. M. Gabbay · 2012
Earlier work this paper cites.
Improving understanding and trust with intelligibility in context-aware applications
B. Y. Lim · 2012
Earlier work this paper cites.
Attributes for classifier feedback
A. Parkash and D. Parikh · 2012
Earlier work this paper cites.
Active learning: Synthesis lectures on artificial intelligence and machine learning
B. Settles · 2012
Earlier work this paper cites.
Simultaneous active learning of classifiers and attributes via relative feedback
A. Biswas and D. Parikh · 2013
Earlier work this paper cites.
Trust in automation
R. R. Hoffman et al · 2013
Earlier work this paper cites.
Power to the people: The role of humans in interactive machine learning
S. Amershi, M. Cakmak, W. B. Knox, and T. Kulesza · 2014
Earlier work this paper cites.
Eliciting good teaching from humans for machine learners
M. Cakmak and A. L. Thomaz · 2014
Earlier work this paper cites.
Classification in the presence of label noise: A survey
B. Frénay and M. Verleysen · 2014
Earlier work this paper cites.
The bayesian case model: A generative approach for case-based reasoning and prototype classification
B. Kim, C. Rudin, and J. A. Shah · 2014
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
K. Simonyan, A. Vedaldi, and A. Zisserman · 2014
Earlier work this paper cites.
Explaining prediction models and individual predictions with feature contributions
E. Štrumbelj and I. Kononenko · 2014
Earlier work this paper cites.
The mind in the machine: Anthropomorphism increases trust in an autonomous vehicle
A. Waytz, J. Heafner, and N. Epley · 2014
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
S. Bach, A. Binder, G. Montavon, F. Klauschen, K.-R. Müller, and W. Samek · 2015
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2015
Earlier work this paper cites.
Principles of explanatory debugging to personalize interactive machine learning
T. Kulesza, M. Burnett, W.-K. Wong, and S. Stumpf · 2015
Earlier work this paper cites.
Explaining Recommendations: Design and Evaluation , pages 353–382
N. Tintarev and J. Masthoff · 2015
Earlier work this paper cites.
Interactive recommender systems: A survey of the state of the art and future research challenges and opportunities
C. He, D. Parra, and K. Verbert · 2016
Earlier work this paper cites.
Interpretable decision sets: A joint framework for description and prediction
H. Lakkaraju, S. H. Bach, and J. Leskovec · 2016
Earlier work this paper cites.
"why should I trust you?": Explaining the predictions of any classifier
M. T. Ribeiro, S. Singh, and C. Guestrin · 2016
Earlier work this paper cites.
Supersparse linear integer models for optimized medical scoring systems
B. Ustun and C. Rudin · 2016
Earlier work this paper cites.
Trust calibration within a human-robot team: Comparing automatically generated explanations
N. Wang et al · 2016
Earlier work this paper cites.
Learning certifiably optimal rule lists for categorical data
E. Angelino, N. Larus-Stone, D. Alabi, M. I. Seltzer, and C. Rudin · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
P. W. Koh and P. Liang · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
S. M. Lundberg and S.-I. Lee · 2017
Earlier work this paper cites.
Snorkel: Rapid training data creation with weak supervision
A. Ratner, S. H. Bach, H. Ehrenberg, J. Fries, S. Wu, and C. Ré · 2017
Earlier work this paper cites.
Right for the right reasons: training differentiable models by constraining their explanations
A. S. Ross, M. C. Hughes, and F. Doshi-Velez · 2017
Earlier work this paper cites.
Grad-CAM: Visual explanations from deep networks via gradient-based localization
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
M. Sundararajan, A. Taly, and Q. Yan · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Counterfactual explanations without opening the black box: Automated decisions and the gdpr
S. Wachter, B. Mittelstadt, and C. Russell · 2017
Earlier work this paper cites.
Peeking inside the black-box: a survey on explainable artificial intelligence (xai)
A. Adadi and M. Berrada · 2018
Earlier work this paper cites.
Sanity checks for saliency maps
J. Adebayo, J. Gilmer, M. Muelly, I. Goodfellow, M. Hardt, and B. Kim · 2018
Earlier work this paper cites.
Towards robust interpretability with self-explaining neural networks
D. Alvarez-Melis and T. S. Jaakkola · 2018
Earlier work this paper cites.
e-snli: natural language inference with natural language explanations
O.-M. Camburu, T. Rocktäschel, T. Lukasiewicz, and P. Blunsom · 2018
Earlier work this paper cites.
Boolean decision rules via column generation
S. Dash, O. Gunluk, and D. Wei · 2018
Earlier work this paper cites.
Explaining explanations: An overview of interpretability of machine learning
L. H. Gilpin, D. Bau, B. Z. Yuan, A. Bajwa, M. Specter, and L. Kagal · 2018
Earlier work this paper cites.
Teaching categories to human learners with visual explanations
O. Mac Aodha, S. Su, Y. Chen, P. Perona, and Y. Yue · 2018
Earlier work this paper cites.
Methods for interpreting and understanding deep neural networks
G. Montavon, W. Samek, and K.-R. Müller · 2018
Earlier work this paper cites.
M. Narayanan, E. Chen, J. He, B. Kim, S. Gershman, and F. Doshi-Velez · 2018
Cited alongside, same era.
Anchors: High-precision model-agnostic explanations
M. T. Ribeiro, S. Singh, and C. Guestrin · 2018
Cited alongside, same era.
A symbolic approach to explaining bayesian network classifiers
A. Shih, A. Choi, and A. Darwiche · 2018
Cited alongside, same era.
Hierarchical interpretations for neural network predictions
C. Singh, W. J. Murdoch, and B. Yu · 2018
Cited alongside, same era.
Is it my looks? or something i said? the impact of explanations, embodiment, and expectations on trust and performance in human-robot teams
N. Wang, D. V. Pynadath, E. Rovira, M. J. Barnes, and S. G. Hill · 2018
Cited alongside, same era.
Machine guides, human supervises: Interactive learning with global explanations
T. Popordanoska, M. Kumar, and S. Teso · 2020
Later among the works it cites.
Interpretations are useful: penalizing explanations to align neural networks with prior knowledge
L. Rieger, C. Singh, W. Murdoch, and B. Yu · 2020
Later among the works it cites.
Protopshare: Prototype sharing for interpretable image classification and similarity discovery
D. Rymarczyk, Ł. Struski, J. Tabor, and B. Zieliński · 2020
Later among the works it cites.
Making deep neural networks right for the right scientific reasons by interacting with their explanations
P. Schramowski, W. Stammer, S. Teso, A. Brugger, F. Herbert, X. Shao, H.-G. Luigs, A.-K. Mahlein, and K. Kersting · 2020
Later among the works it cites.
When explanations lie: Why many modified bp attributions fail
L. Sixt, M. Granz, and T. Landgraf · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Beyond sparsity: Tree regularization of deep models for interpretability
M. Wu, M. Hughes, S. Parbhoo, M. Zazzi, V. Roth, and F. Doshi-Velez · 2018
Cited alongside, same era.
Representer point selection for explaining deep neural networks
C.-K. Yeh, J. S. Kim, I. E. Yen, and P. Ravikumar · 2018
Cited alongside, same era.
Neural-symbolic VQA: disentangling reasoning from vision and language understanding
K. Yi, J. Wu, C. Gan, A. Torralba, P. Kohli, and J. Tenenbaum · 2018
Cited alongside, same era.
Exploring explanation effects on consumers’ trust in online recommender agents
J. Zhang and S. P. Curley · 2018
Cited alongside, same era.
Demystifying black-box models with symbolic metamodels
A. M. Alaa and M. van der Schaar · 2019
Cited alongside, same era.
Where can my career take me? harnessing dialogue for interactive career goal recommendations
Ö. Alkan, E. M. Daly, A. Botea, A. N. Valente, and P. Pedemonte · 2019
Cited alongside, same era.
Counterfactuals in explainable artificial intelligence (XAI): evidence from human reasoning
R. M. J. Byrne · 2019
Cited alongside, same era.
Later among the works it cites.
Towards probabilistic sufficient explanations
E. Wang, P. Khosravi, and G. Van den Broeck · 2020
Later among the works it cites.
Regional tree regularization for interpretability in deep neural networks
M. Wu, S. Parbhoo, M. Hughes, R. Kindle, L. Celi, M. Zazzi, V. Roth, and F. Doshi-Velez · 2020
Later among the works it cites.
Explainable recommendation: A survey and new perspectives
Y. Zhang, X. Chen, et al · 2020
Later among the works it cites.
Beneficial and harmful explanatory machine learning
L. Ai, S. H. Muggleton, C. Hocquette, M. Gromowski, and U. Schmid · 2021
Later among the works it cites.
Irf: A framework for enabling users to interact with recommenders through dialogue
O. Alkan, M. Mattetti, E. M. Daly, A. Botea, I. Vejsbjerg, and B. Knijnenburg · 2021
Later among the works it cites.
Interacting with explanations through critiquing
D. Antognini, C. Musat, and B. Faltings · 2021
Later among the works it cites.
Debiasing concept-based explanations with causal analysis
M. T. Bahadori and D. Heckerman · 2021
Later among the works it cites.
A. J. Barnett, F. R. Schwartz, C. Tao, C. Chen, Y. Ren, J. Y. Lo, and C. Rudin · 2021
Later among the works it cites.
Influence functions in deep learning are fragile
S. Basu, P. Pope, and S. Feizi · 2021
Later among the works it cites.
Explainable machine learning with prior knowledge: An overview
K. Beckh, S. Müller, M. Jakobs, V. Toborek, H. Tan, R. Fischer, P. Welke, S. Houben, and L. von Rueden · 2021
Later among the works it cites.
Principles and practice of explainable machine learning
V. Belle and I. Papantonis · 2021
Later among the works it cites.
Toward a unified framework for debugging gray-box models
A. Bontempelli, F. Giunchiglia, A. Passerini, and S. Teso · 2021
Later among the works it cites.
User driven model adjustment via boolean rule explanations
E. M. Daly, M. Mattetti, Ö. Alkan, and R. Nair · 2021
Later among the works it cites.
Ai for radiographic covid-19 detection selects shortcuts over signal
A. J. DeGrave, J. D. Janizek, and S.-I. Lee · 2021
Later among the works it cites.
On the tractability of SHAP explanations
G. V. den Broeck, A. Lykov, M. Schleich, and D. Suciu · 2021
Later among the works it cites.
Explanation as a process: user-centric construction of multi-level and multi-modal explanations
B. Finzel, D. E. Tafler, S. Scheele, and U. Schmid · 2021
Later among the works it cites.
Causal abstractions of neural networks
A. Geiger, H. Lu, T. Icard, and C. Potts · 2021
Later among the works it cites.
Explainable active learning (xal) toward ai explanations as interfaces for machine teachers
B. Ghai, Q. V. Liao, Y. Zhang, R. Bellamy, and K. Mueller · 2021
Later among the works it cites.
Widening the pipeline in human-guided reinforcement learning with explanation and context-aware data augmentation
L. Guan, M. Verma, S. Guo, R. Zhang, and S. Kambhampati · 2021
Later among the works it cites.
FastIF: Scalable Influence Functions for Efficient Model Interpretation and Debugging
H. Guo, N. Rajani, P. Hase, M. Bansal, and C. Xiong · 2021
Later among the works it cites.
A survey on cost types, interaction schemes, and annotator performance models in selection algorithms for active learning in classification
M. Herde, D. Huseljic, B. Sick, and A. Calma · 2021
Later among the works it cites.
A. Hoffmann, C. Fanconi, R. Rade, and J. Kohler · 2021
Later among the works it cites.
Symbols as a lingua franca for bridging human-ai chasm for explainable and advisable ai systems
S. Kambhampati, S. Sreedharan, M. Verma, Y. Zha, and L. Guan · 2021
Later among the works it cites.
Algorithmic recourse: from counterfactual explanations to interventions
A.-H. Karimi, B. Schölkopf, and I. Valera · 2021
Later among the works it cites.
Sparrow: Semantically coherent prototypes for image classification
S. Kraft, K. Broelemann, A. Theissler, G. Kasneci, G. Esslingen am Neckar, S. H. AG, G. Wiesbaden, and G. Aalen · 2021
Later among the works it cites.
Explanation-based human debugging of nlp models: A survey
P. Lertvittayakumjorn and F. Toni · 2021
Later among the works it cites.
Human-centered explainable ai (xai): From algorithms to user experiences
Q. V. Liao and K. R. Varshney · 2021
Later among the works it cites.
Promises and pitfalls of black-box concept learning models
A. Mahinpei, J. Clark, I. Lage, F. Doshi-Velez, and W. Pan · 2021
Later among the works it cites.
Do concept bottleneck models learn as intended?
A. Margeloiu, M. Ashman, U. Bhatt, Y. Chen, M. Jamnik, and A. Weller · 2021
Later among the works it cites.
Global explanations with decision rules: a co-learning approach
G. Nanfack, P. Temple, and B. Frénay · 2021
Later among the works it cites.
Neural prototype trees for interpretable fine-grained image recognition
M. Nauta, R. van Bree, and C. Seifert · 2021
Later among the works it cites.
Editing a classifier by rewriting its prediction rules
S. Santurkar, D. Tsipras, M. Elango, D. Bau, A. Torralba, and A. Madry · 2021
Later among the works it cites.
Neuro-symbolic artificial intelligence
M. K. Sarker, L. Zhou, A. Eberhart, and P. Hitzler · 2021
Later among the works it cites.
Toward causal representation learning
B. Schölkopf, F. Locatello, S. Bauer, N. R. Ke, N. Kalchbrenner, A. Goyal, and Y. Bengio · 2021
Later among the works it cites.
Glocalx-from local to global explanations of black box ai models
M. Setzu, R. Guidotti, A. Monreale, F. Turini, D. Pedreschi, and F. Giannotti · 2021
Later among the works it cites.
Right for better reasons: Training differentiable models by constraining their influence function
X. Shao, A. Skryagin, P. Schramowski, W. Stammer, and K. Kersting · 2021
Later among the works it cites.
Right for the right concept: Revising neuro-symbolic concepts by interacting with their explanations
W. Stammer, P. Schramowski, and K. Kersting · 2021
Later among the works it cites.
Interactive label cleaning with example-based explanations
S. Teso, A. Bontempelli, F. Giunchiglia, and A. Passerini · 2021
Later among the works it cites.
Notions of explainability and evaluation approaches for explainable artificial intelligence
G. Vilone and L. Longo · 2021
Later among the works it cites.
Saliency is a possible red herring when diagnosing poor generalization
J. D. Viviano, B. Simpson, F. Dutil, Y. Bengio, and J. P. Cohen · 2021
Later among the works it cites.
Informed machine learning - a taxonomy and survey of integrating prior knowledge into learning systems
L. von Rueden, S. Mayer, K. Beckh, B. Georgiev, S. Giesselbach, R. Heese, B. Kirsch, M. Walczak, J. Pfrommer, A. Pick, R. Ramamurthy, J. Garcke, C. Bauckhage, and J. Schuecker · 2021
Later among the works it cites.
Neural-symbolic integration for fairness in ai
B. Wagner and A. d’Avila Garcez · 2021
Later among the works it cites.
Polyjuice: Generating counterfactuals for explaining, evaluating, and improving models
T. Wu, M. T. Ribeiro, J. Heer, and D. S. Weld · 2021
Later among the works it cites.
Causality learning: A new perspective for interpretable machine learning, 2021
G. Xu, T. D. Duong, Q. Li, S. Liu, and X. Wang · 2021
Later among the works it cites.
Refining neural networks with compositional explanations
H. Yao, Y. Chen, Q. Ye, X. Jin, and X. Ren · 2021
Later among the works it cites.
Human-centered concept explanations for neural networks
C. Yeh, B. Kim, and P. Ravikumar · 2021
Later among the works it cites.
Learning from ambiguous demonstrations with self-explanation guided reinforcement learning
Y. Zha, L. Guan, and S. Kambhampati · 2021
Later among the works it cites.
Hildif: Interactive debugging of nli models using influence functions
H. Zylberajch, P. Lertvittayakumjorn, and F. Toni · 2021
Later among the works it cites.
Frote: Feedback rule-driven oversampling for editing models
O. Alkan, D. Wei, M. Mattetti, R. Nair, E. Daly, and D. Saha · 2022
Closest in time.
Concept-level debugging of part-prototype networks
A. Bontempelli, S. Teso, F. Giunchiglia, and A. Passerini · 2022
Closest in time.
G. De Toni, P. Viappiani, B. Lepri, and A. Passerini · 2022
Closest in time.
A typology to explore and guide explanatory interactive machine learning
F. Friedrich, W. Stammer, P. Schramowski, and K. Kersting · 2022
Closest in time.
Building trust in interactive machine learning via user contributed interpretable rules
L. Guo, E. M. Daly, Ö. Alkan, M. Mattetti, O. Cornec, and B. Knijnenburg · 2022
Closest in time.
Explainable deep learning: A field guide for the uninitiated
G. Ras, N. Xie, M. van Gerven, and D. Doran · 2022
Closest in time.
Interpretable machine learning: Fundamental principles and 10 grand challenges
C. Rudin, C. Chen, Z. Chen, H. Huang, L. Semenova, and C. Zhong · 2022
Closest in time.
Right for the right latent factors: Debiasing generative models via disentanglement
X. Shao, K. Stelzner, and K. Kersting · 2022
Closest in time.
Caipi in practice: Towards explainable interactive medical image classification
E. Slany, Y. Ott, S. Scheele, J. Paulus, and U. Schmid · 2022
Closest in time.
Interactive disentanglement: Learning concepts by interacting with their prototype representations
W. Stammer, M. Memmel, P. Schramowski, and K. Kersting · 2022
Closest in time.