Fetching the paper…
Reading the bibliography…
Interpretability techniques in NLP have mainly focused on understanding individual predictions using attention visualization or gradient-based saliency maps over tokens.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
McCoy, R. T.; Pavlick, E.; and Linzen, T. 2019 · 1902
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Zhang, T.; Kishore, V.; Wu, F.; Weinberger, K. Q.; and Artzi, Y. 2019 · 1904
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Reimers, N.; and Gurevych, I. 2019 · 1908
Earlier work this paper cites.
Fine-grained Sentiment Analysis with Faithful Attention
Zhong, R.; Shao, S.; and McKeown, K. 2019 · 1908
Earlier work this paper cites.
Learning to Deceive with Attention-Based Explanations
Pruthi, D.; Gupta, M.; Dhingra, B.; Neubig, G.; and Lipton, Z. C. 2020 · 1909
Earlier work this paper cites.
Reinforced Curriculum Learning on Pre-trained Neural Machine Translation Models
Zhao, M.; Wu, H.; Niu, D.; and Wang, X. 2020 · 2004
Earlier work this paper cites.
The PASCAL recognising textual entailment challenge
Dagan, I.; Glickman, O.; and Magnini, B. 2005 · 2005
Earlier work this paper cites.
Explaining Black Box Predictions and Unveiling Data Artifacts through Influence Functions
Han, X.; Wallace, B. C.; and Tsvetkov, Y. 2020 · 2005
Earlier work this paper cites.
Visualizing data using t-SNE
Maaten, L. v. d.; and Hinton, G. 2008 · 2008
Earlier work this paper cites.
Classification in the presence of label noise: a survey
Frénay, B.; and Verleysen, M. 2013 · 2013
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Bowman, S. R.; Angeli, G.; Potts, C.; and Manning, C. D. 2015 · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S.; and Szegedy, C. 2015 · 2015
Earlier work this paper cites.
Rationalizing Neural Predictions
Lei, T.; Barzilay, R.; and Jaakkola, T. 2016 · 2016
Earlier work this paper cites.
“ Why should i trust you?”
Ribeiro, M. T.; Singh, S.; and Guestrin, C. 2016 · 2016
Cited alongside, same era.
A causal framework for explaining the predictions of black-box sequence-to-sequence models
Alvarez-Melis, D.; and Jaakkola, T. 2017 · 2017
Cited alongside, same era.
Unbounded cache model for online language modeling with open vocabulary
Grave, E.; Cisse, M. M.; and Joulin, A. 2017 · 2017
Cited alongside, same era.
Improving neural language models with a continuous cache
Grave, E.; Joulin, A.; and Usunier, N. 2017 · 2017
Cited alongside, same era.
Understanding black-box predictions via influence functions
Koh, P. W.; and Liang, P. 2017 · 2017
Cited alongside, same era.
Attention is All you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Exs: Explainable search using local model agnostic interpretability
Singh, J.; and Anand, A. 2018 · 2018
Later among the works it cites.
Retrieve and refine: Improved sequence generation models for dialogue
Weston, J.; Dinan, E.; and Miller, A. H. 2018 · 2018
Later among the works it cites.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
Williams, A.; Nangia, N.; and Bowman, S. 2018 · 2018
Later among the works it cites.
Interpretable Neural Predictions with Differentiable Binary Variables
Bastings, J.; Aziz, W.; and Titov, I. 2019 · 2019
Later among the works it cites.
A Game Theoretic Approach to Class-wise Selective Rationalization
Chang, S.; Zhang, Y.; Yu, M.; and Jaakkola, T. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Cited alongside, same era.
Search engine guided neural machine translation
Gu, J.; Wang, Y.; Cho, K.; and Li, V. O. 2018 · 2018
Cited alongside, same era.
Annotation artifacts in natural language inference data
Gururangan, S.; Swayamdipta, S.; Levy, O.; Schwartz, R.; Bowman, S. R.; and Smith, N. A. 2018 · 2018
Cited alongside, same era.
One Sentence One Model for Neural Machine Translation
Li, X.; Zhang, J.; and Zong, C. 2018 · 2018
Cited alongside, same era.
Deep k-nearest neighbors: Towards confident, interpretable and robust deep learning
Papernot, N.; and McDaniel, P. 2018 · 2018
Cited alongside, same era.
Hypothesis only baselines in natural language inference
Poliak, A.; Naradowsky, J.; Haldar, A.; Rudinger, R.; and Van Durme, B. 2018 · 2018
Cited alongside, same era.
Attention is not Explanation
Jain, S.; and Wallace, B. C. 2019 · 2019
Later among the works it cites.
Billion-scale similarity search with GPUs
Johnson, J.; Douze, M.; and Jégou, H. 2019 · 2019
Later among the works it cites.
Is Attention Interpretable?
Serrano, S.; and Smith, N. A. 2019 · 2019
Later among the works it cites.
Attention is not not Explanation
Wiegreffe, S.; and Pinter, Y. 2019 · 2019
Later among the works it cites.
On Identifiability in Transformers
Brunner, G.; Liu, Y.; Pascual, D.; Richter, O.; Ciaramita, M.; and Wattenhofer, R. 2020 · 2020
Closest in time.
Learning The Difference That Makes A Difference With Counterfactually-Augmented Data
Kaushik, D.; Hovy, E.; and Lipton, Z. 2020 · 2020
Closest in time.
Generalization through Memorization: Nearest Neighbor Language Models
Khandelwal, U.; Levy, O.; Jurafsky, D.; Zettlemoyer, L.; and Lewis, M. 2020 · 2020
Closest in time.
Adversarial NLI: A New Benchmark for Natural Language Understanding
Nie, Y.; Williams, A.; Dinan, E.; Bansal, M.; Weston, J.; and Kiela, D. 2020 · 2020
Closest in time.