Fetching the paper…
Reading the bibliography…
For AI systems to garner widespread public acceptance, we must develop methods capable of explaining the decisions of black-box models such as neural networks.
Behavior analysis of NLI models: Uncovering the influence of three factors on robustness
Carmona, V. I. S., Mitchell, J., and Riedel, S. (2018) · 1985
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Learning attitudes and attributes from multi-aspect reviews
McAuley, J. J., Leskovec, J., and Jurafsky, D. (2012) · 2012
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Bowman, S. R., Angeli, G., Potts, C., and Manning, C. D. (2015) · 2015
Earlier work this paper cites.
Molding CNNs for text: Non-linear, non-consecutive convolutions
Lei, T., Barzilay, R., and Jaakkola, T. S. (2015) · 2015
Earlier work this paper cites.
Learning to communicate with deep multi-agent reinforcement learning
Foerster, J. N., Assael, Y. M., de Freitas, N., and Whiteson, S. (2016) · 2016
Earlier work this paper cites.
Rationalizing neural predictions
Lei, T., Barzilay, R., and Jaakkola, T. S. (2016) · 2016
Earlier work this paper cites.
Explaining recurrent neural network predictions in sentiment analysis
Arras, L., Montavon, G., Müller, K.-R., and Samek, W. (2017) · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Koh, P. W. and Liang, P. (2017) · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Lundberg, S. M. and Lee, S.-I. (2017) · 2017
Cited alongside, same era.
Learning important features through propagating activation differences
Shrikumar, A., Greenside, P., and Kundaje, A. (2017) · 2017
Cited alongside, same era.
Sanity checks for saliency maps
Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I. J., Hardt, M., and Kim, B. (2018) · 2018
Cited alongside, same era.
Implementing a robust explanatory bias in a person re-identification network
Bekele, E., Lawson, W. E., Horne, Z., and Khemlani, S. (2018) · 2018
Cited alongside, same era.
e-SNLI: Natural language inference with natural language explanations
Camburu, O., Rocktäschel, T., Lukasiewicz, T., and Blunsom, P. (2018) · 2018
Cited alongside, same era.
What made you do this? Understanding black-box decisions with sufficient input subsets
Annotation artifacts in natural language inference data
Gururangan, S., Swayamdipta, S., Levy, O., Schwartz, R., Bowman, S., and Smith, N. A. (2018) · 2018
Later among the works it cites.
Textual explanations for self-driving vehicles
Kim, J., Rohrbach, A., Darrell, T., Canny, J. F., and Akata, Z. (2018) · 2018
Later among the works it cites.
Defining locality for surrogates in post-hoc interpretablity
Laugel, T., Renard, X., Lesot, M., Marsala, C., and Detyniecki, M. (2018) · 2018
Later among the works it cites.
Multimodal explanations: Justifying decisions and pointing to the evidence
Park, D. H., Hendricks, L. A., Akata, Z., Rohrbach, A., Schiele, B., Darrell, T., and Rohrbach, M. (2018) · 2018
Later among the works it cites.
Model agnostic supervised local explanations
Plumb, G., Molitor, D., and Talwalkar, A. S. (2018) · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Carter, B., Mueller, J., Jain, S., and Gifford, D. K. (2018) · 2018
Cited alongside, same era.
Learning to explain: An information-theoretic perspective on model interpretation
Chen, J., Song, L., Wainwright, M., and Jordan, M. (2018) · 2018
Cited alongside, same era.
Breaking NLI systems with sentences that require simple lexical inferences
Glockner, M., Shwartz, V., and Goldberg, Y. (2018) · 2018
Cited alongside, same era.
Nothing else matters: Model-agnostic explanations by identifying prediction invariance
Ribeiro, M. T., Singh, S., and Guestrin, C. (2016a)
Cited in the paper.
“Why should I trust you?”: Explaining the predictions of any classifier
Ribeiro, M. T., Singh, S., and Guestrin, C. (2016b)
Cited in the paper.
Later among the works it cites.
Evaluating neural network explanation methods using hybrid documents and morphological prediction
Pörner, N., Schütze, H., and Roth, B. (2018) · 2018
Later among the works it cites.
Anchors: High-precision model-agnostic explanations
Ribeiro, M. T., Singh, S., and Guestrin, C. (2018) · 2018
Later among the works it cites.
INVASE: Instance-wise variable selection using neural networks
Yoon, J., Jordon, J., and van der Schaar, M. (2019) · 2019
Closest in time.