Fetching the paper…
Reading the bibliography…
The great success of Transformer-based models benefits from the powerful multi-head self-attention mechanism, which learns token dependencies and encodes contextual information from the input.
Jain, S.; and Wallace, B. C. 2019 · 1902
Earlier work this paper cites.
Are Sixteen Heads Really Better than One?
Michel, P.; Levy, O.; and Neubig, G. 2019 · 1905
Earlier work this paper cites.
What Does BERT Look At? An Analysis of BERT’s Attention
Clark, K.; Khandelwal, U.; Levy, O.; and Manning, C. D. 2019 · 1906
Earlier work this paper cites.
Visualizing and Measuring the Geometry of BERT
Coenen, A.; Reif, E.; Yuan, A.; Kim, B.; Pearce, A.; Viégas, F. B.; and Wattenberg, M. 2019 · 1906
Earlier work this paper cites.
From Balustrades to Pierre Vinken: Looking for Syntax in Transformer Self-Attentions
Marecek, D.; and Rosa, R. 2019 · 1906
Earlier work this paper cites.
Inducing Syntactic Trees from BERT Representations
Rosa, R.; and Marecek, D. 2019 · 1906
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
Long Short-Term Memory
Hochreiter, S.; and Schmidhuber, J. 1997 · 1997
Earlier work this paper cites.
UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training
Bao, H.; Dong, L.; Wei, F.; Wang, W.; Yang, N.; Liu, X.; Wang, Y.; Piao, S.; Gao, J.; Zhou, M.; and Hon, H.-W. 2020 · 2002
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, W. B.; and Brockett, C. 2005 · 2005
Earlier work this paper cites.
The second PASCAL recognising textual entailment challenge
Bar-Haim, R.; Dagan, I.; Dolan, B.; Ferro, L.; and Giampiccolo, D. 2006 · 2006
Earlier work this paper cites.
The PASCAL Recognising Textual Entailment Challenge
Dagan, I.; Glickman, O.; and Magnini, B. 2006 · 2006
Earlier work this paper cites.
InfoXLM: An Information-Theoretic Framework for Cross-Lingual Language Model Pre-Training
Chi, Z.; Dong, L.; Wei, F.; Yang, N.; Singhal, S.; Wang, W.; Song, X.; Mao, X.; Huang, H.; and Zhou, M. 2020b · 2007
Earlier work this paper cites.
The Third PASCAL Recognizing Textual Entailment Challenge
Giampiccolo, D.; Magnini, B.; Dagan, I.; and Dolan, B. 2007 · 2007
Cited alongside, same era.
The Fifth PASCAL Recognizing Textual Entailment Challenge
Bentivogli, L.; Dagan, I.; Dang, H. T.; Giampiccolo, D.; and Magnini, B. 2009 · 2009
Cited alongside, same era.
Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank
Socher, R.; Perelygin, A.; Wu, J.; Chuang, J.; Manning, C. D.; Ng, A.; and Potts, C. 2013 · 2013
Cited alongside, same era.
Neural Machine Translation by Jointly Learning to Align and Translate
Bahdanau, D.; Cho, K.; and Bengio, Y. 2015 · 2015
Cited alongside, same era.
Layer-wise Relevance Propagation for Neural Networks with Local Renormalization Layers
Binder, A.; Montavon, G.; Bach, S.; Müller, K.; and Samek, W. 2016 · 2016
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Later among the works it cites.
Unified Language Model Pre-training for Natural Language Understanding and Generation
Dong, L.; Yang, N.; Wang, W.; Wei, F.; Liu, X.; Wang, Y.; Gao, J.; Zhou, M.; and Hon, H.-W. 2019 · 2019
Later among the works it cites.
A Structural Probe for Finding Syntax in Word Representations
Hewitt, J.; and Manning, C. D. 2019 · 2019
Later among the works it cites.
Revealing the Dark Secrets of BERT
Kovaleva, O.; Romanov, A.; Rogers, A.; and Rumshisky, A. 2019 · 2019
Later among the works it cites.
Is Attention Interpretable?
Serrano, S.; and Smith, N. A. 2019 · 2019
Later among the works it cites.
Universal Adversarial Triggers for Attacking and Analyzing NLP
Wallace, E.; Feng, S.; Kandpal, N.; Gardner, M.; and Singh, S. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Murdoch, W. J.; and Szlam, A. 2017 · 2017
Cited alongside, same era.
Learning Important Features Through Propagating Activation Differences
Shrikumar, A.; Greenside, P.; and Kundaje, A. 2017 · 2017
Cited alongside, same era.
Axiomatic Attribution for Deep Networks
Sundararajan, M.; Taly, A.; and Yan, Q. 2017 · 2017
Cited alongside, same era.
Attention is All you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Interpreting Recurrent and Attention-Based Neural Models: a Case Study on Natural Language Inference
Ghaeini, R.; Fern, X. Z.; and Tadepalli, P. 2018 · 2018
Cited alongside, same era.
Beyond Word Importance: Contextual Decomposition to Extract Interactions from LSTMs
Murdoch, W. J.; Liu, P. J.; and Yu, B. 2018 · 2018
Cited alongside, same era.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
Williams, A.; Nangia, N.; and Bowman, S. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. R. 2019 · 2019
Later among the works it cites.
Attention is not not Explanation
Wiegreffe, S.; and Pinter, Y. 2019 · 2019
Later among the works it cites.
On Identifiability in Transformers
Brunner, G.; Liu, Y.; Pascual, D.; Richter, O.; Ciaramita, M.; and Wattenhofer, R. 2020 · 2020
Closest in time.
Generating Hierarchical Explanations on Text Classification via Feature Interaction Detection
Chen, H.; Zheng, G.; and Ji, Y. 2020 · 2020
Closest in time.
ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators
Clark, K.; Luong, M.-T.; Le, Q. V.; and Manning, C. D. 2020 · 2020
Closest in time.
Unsupervised Cross-lingual Representation Learning at Scale
Conneau, A.; Khandelwal, K.; Goyal, N.; Chaudhary, V.; Wenzek, G.; Guzmán, F.; Grave, E.; Ott, M.; Zettlemoyer, L.; and Stoyanov, V. 2020 · 2020
Closest in time.
Towards Hierarchical Importance Attribution: Explaining Compositional Semantics for Neural Sequence Models
Jin, X.; Wei, Z.; Du, J.; Xue, X.; and Ren, X. 2020 · 2020
Closest in time.