Fetching the paper…
Reading the bibliography…
Transformers are powerful for sequence modeling.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
Efficient Content-Based Sparse Attention with Routing Transformers
Roy, A.; Saffar, M.; Vaswani, A.; and Grangier, D. 2020 · 2003
Earlier work this paper cites.
Bai, H.; Shi, P.; Lin, J.; Tan, L.; Xiong, K.; Gao, W.; Liu, J.; and Li, M. 2020 · 2004
Earlier work this paper cites.
Recipes for building an open-domain chatbot
Roller, S.; Dinan, E.; Goyal, N.; Ju, D.; Williamson, M.; Liu, Y.; Xu, J.; Ott, M.; Shuster, K.; Smith, E. M.; et al. 2020 · 2004
Earlier work this paper cites.
Automatically Constructing a Corpus of Sentential Paraphrases
Dolan, W. B.; and Brockett, C. 2005 · 2005
Earlier work this paper cites.
The Fifth PASCAL Recognizing Textual Entailment Challenge
Bentivogli, L.; Magnini, B.; Dagan, I.; Dang, H. T.; and Giampiccolo, D. 2009 · 2009
Earlier work this paper cites.
Natural Language Processing with Python
Bird, S.; Klein, E.; and Loper, E. 2009 · 2009
Earlier work this paper cites.
Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank
Socher, R.; Perelygin, A.; Wu, J.; Chuang, J.; Manning, C. D.; Ng, A. Y.; and Potts, C. 2013 · 2013
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Bowman, S. R.; Angeli, G.; Potts, C.; and Manning, C. D. 2015 · 2015
Earlier work this paper cites.
Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books
Zhu, Y.; Kiros, R.; Zemel, R. S.; Salakhutdinov, R.; Urtasun, R.; Torralba, A.; and Fidler, S. 2015 · 2015
Earlier work this paper cites.
SQuAD: 100, 000+ Questions for Machine Comprehension of Text
Rajpurkar, P.; Zhang, J.; Lopyrev, K.; and Liang, P. 2016 · 2016
Earlier work this paper cites.
SemEval-2017 Task 1: Semantic Textual Similarity - Multilingual and Cross-lingual Focused Evaluation
Cer, D. M.; Diab, M. T.; Agirre, E.; Lopez-Gazpio, I.; and Specia, L. 2017 · 2017
Earlier work this paper cites.
Improving Neural Language Models with a Continuous Cache
Grave, E.; Joulin, A.; and Usunier, N. 2017 · 2017
Cited alongside, same era.
RACE: Large-scale ReAding Comprehension Dataset From Examinations
Lai, G.; Xie, Q.; Liu, H.; Yang, Y.; and Hovy, E. H. 2017 · 2017
Cited alongside, same era.
Pointer Sentinel Mixture Models
Merity, S.; Xiong, C.; Bradbury, J.; and Socher, R. 2017 · 2017
Cited alongside, same era.
Attention is All you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Dynamic Evaluation of Neural Sequence Models
Krause, B.; Kahembwe, E.; Murray, I.; and Renals, S. 2018 · 2018
Cited alongside, same era.
Improving Language Understanding by Generative Pre-Training
Radford, A. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2019 · 2019
Later among the works it cites.
Attribute-aware Sequence Network for Review Summarization
Li, J.; Wang, X.; Yin, D.; and Zong, C. 2019 · 2019
Later among the works it cites.
Text Summarization with Pretrained Encoders
Liu, Y.; and Lapata, M. 2019 · 2019
Later among the works it cites.
Improving Question Answering with External Knowledge
Pan, X.; Sun, K.; Yu, D.; Chen, J.; Ji, H.; Cardie, C.; and Yu, D. 2019 · 2019
Later among the works it cites.
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reimers, N.; and Gurevych, I. 2019 · 2019
Later among the works it cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. R. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fast Parametric Learning with Activation Memorization
Rae, J. W.; Dyer, C.; Dayan, P.; and Lillicrap, T. P. 2018 · 2018
Cited alongside, same era.
Know What You Don’t Know: Unanswerable Questions for SQuAD
Rajpurkar, P.; Jia, R.; and Liang, P. 2018 · 2018
Cited alongside, same era.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
Williams, A.; Nangia, N.; and Bowman, S. R. 2018 · 2018
Cited alongside, same era.
Character-Level Language Modeling with Deeper Self-Attention
Al-Rfou, R.; Choe, D.; Constant, N.; Guo, M.; and Jones, L. 2019 · 2019
Cited alongside, same era.
Adaptive Input Representations for Neural Language Modeling
Baevski, A.; and Auli, M. 2019 · 2019
Cited alongside, same era.
Transformer-XL: Attentive Language Models beyond a Fixed-Length Context
Dai, Z.; Yang, Z.; Yang, Y.; Carbonell, J. G.; Le, Q. V.; and Salakhutdinov, R. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Neural network acceptability judgments
Warstadt, A.; Singh, A.; and Bowman, S. R. 2019 · 2019
Later among the works it cites.
XLNet: Generalized Autoregressive Pretraining for Language Understanding
Yang, Z.; Dai, Z.; Yang, Y.; Carbonell, J. G.; Salakhutdinov, R.; and Le, Q. V. 2019 · 2019
Later among the works it cites.
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Lan, Z.; Chen, M.; Goodman, S.; Gimpel, K.; Sharma, P.; and Soricut, R. 2020 · 2020
Closest in time.
Compressive Transformers for Long-Range Sequence Modelling
Rae, J. W.; Potapenko, A.; Jayakumar, S. M.; Hillier, C.; and Lillicrap, T. P. 2020 · 2020
Closest in time.
StructBERT: Incorporating Language Structures into Pre-training for Deep Language Understanding
Wang, W.; Bi, B.; Yan, M.; Wu, C.; Bao, Z.; Peng, L.; and Si, L. 2020 · 2020
Closest in time.