Fetching the paper…
Reading the bibliography…
The ability to identify influential training examples enables us to debug training data and explain model behavior.
Residuals and influence in regression
Cook, R. D. and Weisberg, S · 1982
Earlier work this paper cites.
Term-weighting approaches in automatic text retrieval
Salton, G. and Buckley, C · 1988
Earlier work this paper cites.
How to explain individual classification decisions
Baehrens, D., Schroeter, T., Harmeling, S., Kawanabe, M., Hansen, K., and MÞller, K.-R · 2010
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E · 2011
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Simonyan, K., Vedaldi, A., and Zisserman, A · 2013
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Zeiler, M. D. and Fergus, R · 2014
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Bach, S., Binder, A., Montavon, G., Klauschen, F., Müller, K.-R., and Samek, W · 2015
Earlier work this paper cites.
Ag corpus of news articles
Gulli, A · 2015
Earlier work this paper cites.
Falling rule lists
Wang, F. and Rudin, C · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Zhang, X., Zhao, J., and LeCun, Y · 2015
Earlier work this paper cites.
Why should i trust you?: Explaining the predictions of any classifier
Ribeiro, M. T., Singh, S., and Guestrin, C · 2016
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Koh, P. W. and Liang, P · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Lundberg, S. M. and Lee, S.-I · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences
Shrikumar, A., Greenside, P., and Kundaje, A · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Sundararajan, M., Taly, A., and Yan, Q · 2017
Earlier work this paper cites.
Counterfactual explanations without opening the black box: Automated decisions and the gdpr
Wachter, S., Mittelstadt, B. D., and Russell, C · 2017
Earlier work this paper cites.
A unified view of gradient-based attribution methods for deep neural networks
Ancona, M., Ceolini, E., Öztireli, C., and Gross, M · 2018
Earlier work this paper cites.
Explanations based on the missing: Towards contrastive explanations with pertinent negatives
Dhurandhar, A., Chen, P.-Y., Luss, R., Tu, C.-C., Ting, P., Shanmugam, K., and Das, P · 2018
Cited alongside, same era.
Grounding visual explanations
Hendricks, L. A., Hu, R., Darrell, T., and Akata, Z · 2018
Cited alongside, same era.
Toxic comment classification challenge: Identify and classify toxic online comments
Kaggle.com · 2018
Cited alongside, same era.
Interpreting black box predictions using fisher kernels
Khanna, R., Kim, B., Ghosh, J., and Koyejo, O · 2018
Cited alongside, same era.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., Viegas, F., et al · 2018
Cited alongside, same era.
Exploring principled visualizations for deep network attributions
Sundararajan, M., Xu, J., Taly, A., Sayres, R., and Najmi, A · 2019
Later among the works it cites.
Allennlp interpret: A framework for explaining predictions of nlp models
Wallace, E., Tuyls, J., Wang, J., Subramanian, S., Gardner, M., and Singh, S · 2019
Later among the works it cites.
On the (in)fidelity and sensitivity of explanations
Yeh, C.-K., Hsieh, C.-Y., Suggala, A. S., Inouye, D. I., and Ravikumar, P · 2019
Later among the works it cites.
Relatif: Identifying explanatory training samples via relative influence
Barshan, E., Brunet, M.-E., and Dziugaite, G. K · 2020
Later among the works it cites.
Fortifying toxic speech detectors against disguised toxicity
Han, X. and Tsvetkov, Y · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Petsiuk, V., Das, A., and Saenko, K · 2018
Cited alongside, same era.
Contrastive Explanations with Local Foil Trees
van der Waa, J., Robeer, M., van Diggelen, J., Brinkhuis, M., and Neerincx, M · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A., Nangia, N., and Bowman, S · 2018
Cited alongside, same era.
Representer point selection for explaining deep neural networks
Yeh, C.-K., Kim, J., Yen, I. E.-H., and Ravikumar, P. K · 2018
Cited alongside, same era.
Interpretable basis decomposition for visual explanation
Zhou, B., Sun, Y., Bau, D., and Torralba, A · 2018
Cited alongside, same era.
This looks like that: deep learning for interpretable image recognition
Chen, C., Li, O., Tao, D., Barnett, A., Rudin, C., and Su, J. K · 2019
Cited alongside, same era.
Counterfactual visual explanations
Goyal, Y., Wu, Z., Ernst, J., Batra, D., Parikh, D., and Lee, S · 2019
Cited alongside, same era.
Han, X., Wallace, B. C., and Tsvetkov, Y · 2020
Later among the works it cites.
Deberta: Decoding-enhanced bert with disentangled attention
He, P., Liu, X., Gao, J., and Chen, W · 2020
Later among the works it cites.
On the sentence embeddings from pre-trained language models
Li, B., Zhou, H., He, J., Wang, M., Yang, Y., and Li, L · 2020
Later among the works it cites.
The penalty imposed by ablated data augmentation
Liu, F., Najmi, A., and Sundararajan, M · 2020
Later among the works it cites.
Estimating training data influence by tracing gradient descent
Pruthi, G., Liu, F., Kale, S., and Sundararajan, M · 2020
Later among the works it cites.
Han, X. and Tsvetkov, Y · 2021
Later among the works it cites.
Evaluation of similarity-based explanations
Hanawa, K., Yokoi, S., Hara, S., and Inui, K · 2021
Later among the works it cites.
Guided integrated gradients: An adaptive path method for removing noise
Kapishnikov, A., Venugopalan, S., Avci, B., Wedin, B., Terry, M., and Bolukbasi, T · 2021
Later among the works it cites.
Combining feature and instance attribution to detect artifacts
Pezeshkpour, P., Jain, S., Singh, S., and Wallace, B. C · 2021
Later among the works it cites.
Revisiting methods for finding influential examples
Søgaard, A. et al · 2021
Later among the works it cites.
Representer point selection via local jacobian expansion for post-hoc classifier explanation of deep neural networks and ensemble models
Sui, Y., Wu, G., and Sanner, S · 2021
Later among the works it cites.