Fetching the paper…
Reading the bibliography…
An extractive rationale explains a language model's (LM's) prediction on a given task instance by highlighting the text inputs that most influenced the prediction.
Measuring nominal scale agreement among many raters
Fleiss, J. L · 1971
Earlier work this paper cites.
Modeling annotators: A generative approach to learning from annotator rationales
Zaidan, O. and Eisner, J · 2008
Earlier work this paper cites.
Hidden factors and hidden topics: understanding rating dimensions with review text
McAuley, J. and Leskovec, J · 2013
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Simonyan, K., Vedaldi, A., and Zisserman, A · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C · 2013
Earlier work this paper cites.
Extraction of salient sentences from labelled documents
Denil, M., Demiraj, A., and De Freitas, N · 2014
Earlier work this paper cites.
Visualizing and understanding neural models in nlp
Li, J., Chen, X., Hovy, E., and Jurafsky, D · 2015
Earlier work this paper cites.
Character-level Convolutional Networks for Text Classification
Zhang, X., Zhao, J., and LeCun, Y · 2015
Earlier work this paper cites.
Rationalizing neural predictions
Lei, T., Barzilay, R., and Jaakkola, T · 2016
Earlier work this paper cites.
Understanding neural networks through representation erasure
Li, J., Monroe, W., and Jurafsky, D · 2016
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Doshi-Velez, F. and Kim, B · 2017
Earlier work this paper cites.
Representation of linguistic form and function in recurrent neural networks
Kádár, A., Chrupała, G., and Alishahi, A · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Lundberg, S. M. and Lee, S.-I · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences
Shrikumar, A., Greenside, P., and Kundaje, A · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Sundararajan, M., Taly, A., and Yan, Q · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
e-snli: Natural language inference with natural language explanations
Camburu, O.-M., Rocktäschel, T., Lukasiewicz, T., and Blunsom, P · 2018
Earlier work this paper cites.
Hate speech dataset from a white supremacy forum
de Gibert, O., Perez, N., García-Pablos, A., and Cuadros, M · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
A benchmark for interpretability methods in deep neural networks
Hooker, S., Erhan, D., Kindermans, P.-J., and Kim, B · 2018
Earlier work this paper cites.
Looking beyond the surface: A challenge set for reading comprehension over multiple sentences
Khashabi, D., Chaturvedi, S., Roth, M., Upadhyay, S., and Roth, D · 2018
Cited alongside, same era.
The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery
Lipton, Z. C · 2018
Cited alongside, same era.
Evaluating neural network explanation methods using hybrid documents and morphological agreement
Poerner, N., Roth, B., and Schütze, H · 2018
Cited alongside, same era.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Talmor, A., Herzig, J., Lourie, N., and Berant, J · 2018
Cited alongside, same era.
Semeval-2018 task 3: Irony detection in english tweets
Van Hee, C., Lefever, E., and Hoste, V · 2018
Cited alongside, same era.
Learning to faithfully rationalize by construction
Jain, S., Wiegreffe, S., Pinter, Y., and Wallace, B. C · 2020
Later among the works it cites.
Captum: A unified and generic model interpretability library for pytorch
Kokhlikyan, N., Miglani, V., Martin, M., Wang, E., Alsallakh, B., Reynolds, J., Melnikov, A., Kliushkina, N., Araya, C., Yan, S., et al · 2020
Later among the works it cites.
Nile: Natural language inference with faithful natural language explanations
Kumar, S. and Talukdar, P · 2020
Later among the works it cites.
Fid-ex: Improving sequence-to-sequence models for extractive rationale generation
Lakhotia, K., Paranjape, B., Ghoshal, A., Yih, W.-t., Mehdad, Y., and Iyer, S · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Interpretable neural predictions with differentiable binary variables
Bastings, J., Aziz, W., and Titov, I · 2019
Cited alongside, same era.
Eraser: A benchmark to evaluate rationalized nlp models
DeYoung, J., Jain, S., Rajani, N. F., Lehman, E., Xiong, C., Socher, R., and Wallace, B. C · 2019
Cited alongside, same era.
PyTorch Lightning, 3 2019
Falcon, W. and The PyTorch Lightning team · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Cited alongside, same era.
Explain yourself! leveraging language models for commonsense reasoning
Rajani, N. F., McCann, B., Xiong, C., and Socher, R · 2019
Cited alongside, same era.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Rudin, C · 2019
Cited alongside, same era.
Narang, S., Raffel, C., Lee, K., Roberts, A., Fiedel, N., and Malkan, K · 2020
Later among the works it cites.
An information bottleneck approach for controlling conciseness in rationale extraction
Paranjape, B., Joshi, M., Thickstun, J., Hajishirzi, H., and Zettlemoyer, L · 2020
Later among the works it cites.
Evaluating explanations: How much do explanations from the teacher aid students?
Pruthi, D., Dhingra, B., Soares, L. B., Collins, M., Lipton, Z. C., Neubig, G., and Cohen, W. W · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Zaheer, M., Guruganesh, G., Dubey, K. A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., et al · 2020
Later among the works it cites.
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S · 2021
Closest in time.
Self-training with few-shot rationalization: Teacher explanations aid student in few-shot nlu
Bhat, M. M., Sordoni, A., and Mukherjee, S · 2021
Closest in time.
Salkg: Learning from knowledge graph explanations for commonsense reasoning
Chan, A., Xu, J., Long, B., Sanyal, S., Gupta, T., and Ren, X · 2021
Closest in time.
Improving deep learning interpretability by saliency guided training
Ismail, A. A., Corrada Bravo, H., and Feizi, S · 2021
Closest in time.
Local interpretations for explainable natural language processing: A survey
Luo, S., Ivison, H., Han, C., and Poon, J · 2021
Closest in time.
Deep learning–based text classification: A comprehensive review
Minaee, S., Kalchbrenner, N., Cambria, E., Nikzad, N., Chenaghlu, M., and Gao, J · 2021
Closest in time.
Discretized integrated gradients for explaining language models
Sanyal, S. and Ren, X · 2021
Closest in time.
Efficient explanations from empirical explainers
Schwarzenberg, R., Feldhus, N., and Möller, S · 2021
Closest in time.
Learning to explain: Generating stable explanations fast
Situ, X., Zukerman, I., Paris, C., Maruf, S., and Haffari, G · 2021
Closest in time.
Crossfit: A few-shot learning challenge for cross-task generalization in nlp
Ye, Q., Lin, B. Y., and Ren, X · 2021
Closest in time.
Understanding interlocking dynamics of cooperative rationalization
Yu, M., Zhang, Y., Chang, S., and Jaakkola, T · 2021
Closest in time.