Fetching the paper…
Reading the bibliography…
Self attention networks (SANs) have been widely utilized in recent NLP studies.
Bidirectional recurrent neural networks
Schuster, M.; and Paliwal, K. K. 1997 · 1997
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R.; Perelygin, A.; Wu, J.; Chuang, J.; Manning, C. D.; Ng, A. Y.; and Potts, C. 2013 · 2013
Earlier work this paper cites.
A fast and accurate dependency parser using neural networks
Chen, D.; and Manning, C. D. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Bowman, S.; Angeli, G.; Potts, C.; and Manning, C. D. 2015 · 2015
Earlier work this paper cites.
When Are Tree Structures Necessary for Deep Learning of Representations?
Li, J.; Luong, M.-T.; Jurafsky, D.; and Hovy, E. 2015 · 2015
Earlier work this paper cites.
Improved Semantic Representations From Tree-Structured Long Short-Term Memory Networks
Tai, K. S.; Socher, R.; and Manning, C. D. 2015 · 2015
Earlier work this paper cites.
A Fast Unified Model for Parsing and Sentence Understanding
Bowman, S.; Gauthier, J.; Rastogi, A.; Gupta, R.; Manning, C. D.; and Potts, C. 2016 · 2016
Earlier work this paper cites.
Natural Language Inference by Tree-Based Convolution and Heuristic Matching
Mou, L.; Men, R.; Li, G.; Xu, Y.; Zhang, L.; Yan, R.; and Jin, Z. 2016 · 2016
Earlier work this paper cites.
Distance-based self-attention network for natural language inference
Im, J.; and Cho, S. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A.; Nangia, N.; and Bowman, S. R. 2017 · 2017
Cited alongside, same era.
Learning to compose task-specific tree structures
Choi, J.; Yoo, K. M.; and Lee, S.-g. 2018 · 2018
Cited alongside, same era.
Self-Attention with Relative Position Representations
Shaw, P.; Uszkoreit, J.; and Vaswani, A. 2018 · 2018
Cited alongside, same era.
Phrase-level self-attention networks for universal sentence encoding
Wu, W.; Wang, H.; Liu, T.; and Ma, S. 2018 · 2018
Cited alongside, same era.
Dynamic self-attention: Computing attention over words dynamically for sentence embedding
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Later among the works it cites.
Gaussian transformer: a lightweight approach for natural language inference
Guo, M.; Zhang, Y.; and Liu, T. 2019 · 2019
Later among the works it cites.
Star-Transformer
Guo, Q.; Qiu, X.; Liu, P.; Shao, Y.; Xue, X.; and Zhang, Z. 2019 · 2019
Later among the works it cites.
A structural probe for finding syntax in word representations
Hewitt, J.; and Manning, C. D. 2019 · 2019
Later among the works it cites.
Sentence embeddings in NLI with iterative refinement encoders
Talman, A.; Yli-Jyrä, A.; and Tiedemann, J. 2019 · 2019
Later among the works it cites.
Self-Attention with Structural Position Representations
Wang, X.; Tu, Z.; Wang, L.; and Shi, S. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yoon, D.; Lee, D.; and Lee, S. 2018 · 2018
Cited alongside, same era.
Qanet: Combining local convolution with global self-attention for reading comprehension
Yu, A. W.; Dohan, D.; Luong, M.-T.; Zhao, R.; Chen, K.; Norouzi, M.; and Le, Q. V. 2018 · 2018
Cited alongside, same era.
Transformer-XL: Attentive Language Models beyond a Fixed-Length Context
Dai, Z.; Yang, Z.; Yang, Y.; Carbonell, J. G.; Le, Q.; and Salakhutdinov, R. 2019 · 2019
Cited alongside, same era.
Disan: Directional self-attention network for rnn/cnn-free language understanding
Shen, T.; Zhou, T.; Long, G.; Jiang, J.; Pan, S.; and Zhang, C. 2018a
Cited in the paper.
Reinforced self-attention network: a hybrid of hard and soft attention for sequence modeling
Shen, T.; Zhou, T.; Long, G.; Jiang, J.; Wang, S.; and Zhang, C. 2018b
Cited in the paper.
Tree Transformer: Integrating Tree Structures into Self-Attention
Wang, Y.; Lee, H.-Y.; and Chen, Y.-N. 2019 · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Yang, Z.; Dai, Z.; Yang, Y.; Carbonell, J.; Salakhutdinov, R. R.; and Le, Q. V. 2019 · 2019
Later among the works it cites.