Gender shades: Intersectional accuracy disparities in commercial gender classification
Buolamwini, J. and Gebru, T. (2018) · 2018
Later among the works it cites.
e-SNLI: Natural language inference with natural language explanations
Camburu, O.-M., Rocktäschel, T., Lukasiewicz, T., and Blunsom, P. (2018) · 2018
Later among the works it cites.
Boosting adversarial attacks with momentum
Dong, Y., Liao, F., Pang, T., Su, H., Zhu, J., Hu, X., and Li, J. (2018) · 2018
Later among the works it cites.
Empirical risk minimization under fairness constraints
Donini, M., Oneto, L., Ben-David, S., Shawe-Taylor, J. S., and Pontil, M. (2018) · 2018
Later among the works it cites.
Considerations for evaluation and generalization in interpretable machine learning
Doshi-Velez, F. and Kim, B. (2018) · 2018
Later among the works it cites.
Explainable artificial intelligence: A survey
Došilović, F. K., Brčić, M., and Hlupić, N. (2018) · 2018
Later among the works it cites.
Decoupled classifiers for group-fair and efficient machine learning
Dwork, C., Immorlica, N., Kalai, A. T., and Leiserson, M. (2018) · 2018
Later among the works it cites.
Robust physical-world attacks on deep learning visual classification
Eykholt, K., Evtimov, I., Fernandes, E., Li, B., Rahmati, A., Xiao, C., Prakash, A., Kohno, T., and Song, D. (2018) · 2018
Later among the works it cites.
Neural stethoscopes: Unifying analytic, auxiliary and adversarial network probing
Original
Fuchs, F. B., Groth, O., Kosoriek, A. R., Bewley, A., Wulfmeier, M., Vedaldi, A., and Posner, I. (2018) · 2018
Later among the works it cites.
Explaining explanations: An overview of interpretability of machine learning
Gilpin, L. H., Bau, D., Yuan, B. Z., Bajwa, A., Specter, M., and Kagal, L. (2018) · 2018
Later among the works it cites.
Do semantic parts emerge in convolutional neural networks?
Gonzalez-Garcia, A., Modolo, D., and Ferrari, V. (2018) · 2018
Later among the works it cites.
Shapestacks: Learning vision-based physical intuition for generalised object stacking
Groth, O., Fuchs, F. B., Posner, I., and Vedaldi, A. (2018) · 2018
Later among the works it cites.
A survey of methods for explaining black box models
Guidotti, R., Monreale, A., Ruggieri, S., Turini, F., Giannotti, F., and Pedreschi, D. (2018) · 2018
Later among the works it cites.
Causal learning and explanation of deep neural networks via autoencoded activations
Original
Harradon, M., Druce, J., and Ruttenberg, B. (2018) · 2018
Later among the works it cites.
Effective attention modeling for aspect-level sentiment classification
He, R., Lee, W. S., Ng, H. T., and Dahlmeier, D. (2018) · 2018
Later among the works it cites.
Fairness behind a veil of ignorance: A welfare analysis for automated decision making
Heidari, H., Ferrari, C., Gummadi, K., and Krause, A. (2018) · 2018
Later among the works it cites.
Uncertainty-aware attention for reliable interpretation and prediction
Heo, J., Lee, H. B., Kim, S., Lee, J., Kim, K. J., Yang, E., and Hwang, S. J. (2018) · 2018
Later among the works it cites.
Transparency and explanation in deep reinforcement learning neural networks
Iyer, R., Li, Y., Li, H., Lewis, M., Sundar, R., and Sycara, K. (2018) · 2018
Later among the works it cites.
To trust or not to trust a classifier
Jiang, H., Kim, B., Guan, M., and Gupta, M. (2018) · 2018
Later among the works it cites.
Model assertions for debugging machine learning
Kang, D., Raghavan, D., Bailis, P., and Zaharia, M. (2018) · 2018
Later among the works it cites.
Preventing fairness gerrymandering: Auditing and learning for subgroup fairness
Kearns, M., Neel, S., Roth, A., and Wu, Z. S. (2018) · 2018
Later among the works it cites.
Learning how to explain neural networks: Patternnet and Patternattribution
Kindermans, P.-J., Schütt, K., Alber, M., Müller, K.-R., Erhan, D., Kim, B., and Dähne, S. (2018) · 2018
Later among the works it cites.
Importance of self-attention for sentiment analysis
Letarte, G., Paradis, F., Giguère, P., and Laviolette, F. (2018) · 2018
Later among the works it cites.
The mythos of model interpretability
Lipton, Z. C. (2018) · 2018
Later among the works it cites.
Improving the interpretability of deep neural networks with knowledge distillation
Liu, X., Wang, X., and Matwin, S. (2018b) · 2018
Later among the works it cites.
Consistent individualized feature attribution for tree ensembles
Original
Lundberg, S. M., Erion, G. G., and Lee, S.-I. (2018) · 2018
Later among the works it cites.
Transparency by design: Closing the gap between performance and interpretability in visual reasoning
Mascharka, D., Tran, P., Soklaski, R., and Majumdar, A. (2018) · 2018
Later among the works it cites.
Deep face recognition: A survey
Masi, I., Wu, Y., Hassner, T., and Natarajan, P. (2018) · 2018
Later among the works it cites.
Towards robust interpretability with self-explaining neural networks
Melis, D. A. and Jaakkola, T. (2018) · 2018
Later among the works it cites.
The cost of fairness in binary classification
Menon, A. K. and Williamson, R. C. (2018) · 2018
Later among the works it cites.
Contrastive explanation: A structural-model approach
Original
Miller, T. (2018) · 2018
Later among the works it cites.
Methods for interpreting and understanding deep neural networks
Montavon, G., Samek, W., and Müller, K.-R. (2018) · 2018
Later among the works it cites.
Beyond word importance: Contextual decomposition to extract interactions from LSTMs
Original
Murdoch, W. J., Liu, P. J., and Yu, B. (2018) · 2018
Later among the works it cites.
An interpretable machine learning model for accurate prediction of sepsis in the ICU
Nemati, S., Holder, A., Razmi, F., Stanley, M. D., Clifford, G. D., and Buchman, T. G. (2018) · 2018
Later among the works it cites.
A theoretical explanation for perplexing behaviors of backpropagation-based visualizations
Nie, W., Zhang, Y., and Patel, A. (2018) · 2018
Later among the works it cites.
The building blocks of interpretability
Olah, C., Satyanarayan, A., Johnson, I., Carter, S., Schubert, L., Ye, K., and Mordvintsev, A. (2018) · 2018
Later among the works it cites.
Multimodal explanations: Justifying decisions and pointing to the evidence
Park, D. H., Hendricks, L. A., Akata, Z., Rohrbach, A., Schiele, B., Darrell, T., and Rohrbach, M. (2018) · 2018
Later among the works it cites.
Rise: Randomized input sampling for explanation of black-box models
Original
Petsiuk, V., Das, A., and Saenko, K. (2018) · 2018
Later among the works it cites.
Deep learning for chest radiograph diagnosis: A retrospective comparison of the CheXNeXt algorithm to practicing radiologists
Rajpurkar, P., Irvin, J., Ball, R. L., Zhu, K., Yang, B., Mehta, H., Duan, T., Ding, D., Bagul, A., Langlotz, C. P., et al. (2018) · 2018
Later among the works it cites.
Explanation methods in deep learning: Users, values, concerns and challenges
Ras, G., van Gerven, M., and Haselager, P. (2018) · 2018
Later among the works it cites.
Anchors: High-precision model-agnostic explanations
Ribeiro, M. T., Singh, S., and Guestrin, C. (2018) · 2018
Later among the works it cites.
Perturbation-based explanations of prediction models
Robnik-Šikonja, M. and Bohanec, M. (2018) · 2018
Later among the works it cites.
Defense-gan: Protecting classifiers against adversarial attacks using generative models
Original
Samangouei, P., Kabkab, M., and Chellappa, R. (2018) · 2018
Later among the works it cites.
Learning global additive explanations for neural nets using model distillation
Original
Tan, S., Caruana, R., Hooker, G., Koch, P., and Gordo, A. (2018) · 2018
Later among the works it cites.
Survey on virtual assistant: Google Assistant, Siri, Cortana, Alexa
Tulshan, A. S. and Dhage, S. N. (2018) · 2018
Later among the works it cites.
Interpreting cnn knowledge via an explanatory graph
Zhang, Q., Cao, R., Shi, F., Wu, Y. N., and Zhu, S.-C. (2018) · 2018
Later among the works it cites.
Visual interpretability for deep learning: A survey
Zhang, Q.-s. and Zhu, S.-C. (2018) · 2018
Later among the works it cites.
Activation atlas
Carter, S., Armstrong, Z., Schubert, L., Johnson, I., and Olah, C. (2019) · 2019
Later among the works it cites.
Machine learning interpretability: A survey on methods and metrics
Carvalho, D. V., Pereira, E. M., and Cardoso, J. S. (2019) · 2019
Later among the works it cites.
This looks like that: Deep learning for interpretable image recognition
Chen, C., Li, O., Tao, D., Barnett, A., Rudin, C., and Su, J. K. (2019) · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2019) · 2019
Later among the works it cites.
Obtaining fairness using optimal transport theory
Gordaliza, P., Del Barrio, E., Fabrice, G., and Loubes, J.-M. (2019) · 2019
Later among the works it cites.
TED: Teaching AI to explain its decisions
Hind, M., Wei, D., Campbell, M., Codella, N. C., Dhurandhar, A., Mojsilović, A., Natesan Ramamurthy, K., and Varshney, K. R. (2019) · 2019
Later among the works it cites.
A benchmark for interpretability methods in deep neural networks
Hooker, S., Erhan, D., Kindermans, P.-J., and Kim, B. (2019) · 2019
Later among the works it cites.
Attention is not explanation
Jain, S. and Wallace, B. C. (2019) · 2019
Later among the works it cites.
The (un)reliability of saliency methods
Kindermans, P.-J., Hooker, S., Adebayo, J., Alber, M., Schütt, K. T., Dähne, S., Erhan, D., and Kim, B. (2019) · 2019
Later among the works it cites.
Towards explainable NLP: A generative explanation framework for text classification
Liu, H., Yin, Q., and Wang, W. Y. (2019) · 2019
Later among the works it cites.
Explanation in artificial intelligence: Insights from the social sciences
Miller, T. (2019) · 2019
Later among the works it cites.
PyTorch CNN visualizations
Ozbulak, U. (2019) · 2019
Later among the works it cites.
Can you explain that? Lucid explanations help human-AI collaborative image retrieval
Ray, A., Yao, Y., Kumar, R., Divakaran, A., and Burachas, G. (2019) · 2019
Later among the works it cites.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Rudin, C. (2019) · 2019
Later among the works it cites.
One pixel attack for fooling deep neural networks
Su, J., Vargas, D. V., and Sakurai, K. (2019) · 2019
Later among the works it cites.
Attention is not not explanation
Wiegreffe, S. and Pinter, Y. (2019) · 2019
Later among the works it cites.
A formal approach to explainability
Wolf, L., Galanti, T., and Hazan, T. (2019) · 2019
Later among the works it cites.
Adversarial examples: Attacks and defenses for deep learning
Yuan, X., He, P., Zhu, Q., and Li, X. (2019) · 2019
Later among the works it cites.
From recognition to cognition: Visual commonsense reasoning
Zellers, R., Bisk, Y., Farhadi, A., and Choi, Y. (2019) · 2019
Later among the works it cites.
Explainable artificial intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI
Arrieta, A. B., Díaz-Rodríguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., García, S., Gil-López, S., Molina, D., Benjamins, R., et al. (2020) · 2020
Closest in time.
Statistical mechanics of deep learning
Bahri, Y., Kadmon, J., Pennington, J., Schoenholz, S. S., Sohl-Dickstein, J., and Ganguli, S. (2020) · 2020
Closest in time.
A survey of deep learning techniques for autonomous driving
Grigorescu, S., Trasnea, B., Cocias, T., and Macesanu, G. (2020) · 2020
Closest in time.
Causal shapley values: Exploiting causal knowledge to explain individual predictions of complex models
Heskes, T., Sijben, E., Bucur, I. G., and Claassen, T. (2020) · 2020
Closest in time.
Learning with interpretable structure from gated RNN
Hou, B.-J. and Zhou, Z.-H. (2020) · 2020
Closest in time.
Interpretable Machine Learning
Molnar, C. (2020) · 2020
Closest in time.
How can I choose an explainer? An application-grounded evaluation of post-hoc explanations
Original
Jesus, S., Belém, C., Balayan, V., Bento, J., Saleiro, P., Bizarro, P., and Gama, J. (2021) · 2021
Closest in time.
Adaptive deconvolutional networks for mid and high level feature learning
Zeiler, M. D., Taylor, G. W., and Fergus, R. (2011) · 2025
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., and Bengio, Y. (2015) · 2057
Closest in time.