Fetching the paper…
Reading the bibliography…
Pre-training language models (LMs) on large-scale unlabeled text data makes the model much easier to achieve exceptional downstream performance than their counterparts directly trained on the downstream tasks.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
HuggingFace’s Transformers: State-of-the-art Natural Language Processing
Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; Davison, J.; Shleifer, S.; von Platen, P.; Ma, C.; Jernite, Y.; Plu, J.; Xu, C.; Scao, T. L.; Gugger, S.; Drame, M.; Lhoest, Q.; and Rush, A. M. 2019 · 1910
Earlier work this paper cites.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
Dodge, J.; Ilharco, G.; Schwartz, R.; Farhadi, A.; Hajishirzi, H.; and Smith, N. 2020 · 2002
Earlier work this paper cites.
Automatically Constructing a Corpus of Sentential Paraphrases
Dolan, W. B.; and Brockett, C. 2005 · 2005
Earlier work this paper cites.
A Mathematical Exploration of Why Language Models Help Solve Downstream Tasks
Saunshi, N.; Malladi, S.; and Arora, S. 2020 · 2010
Earlier work this paper cites.
Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank
Socher, R.; Perelygin, A.; Wu, J.; Chuang, J.; Manning, C. D.; Ng, A.; and Potts, C. 2013 · 2013
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Rajpurkar, P.; Zhang, J.; Lopyrev, K.; and Liang, P. 2016 · 2016
Earlier work this paper cites.
SemEval-2017 Task 1: Semantic Textual Similarity Multilingual and Crosslingual Focused Evaluation
Cer, D.; Diab, M.; Agirre, E.; Lopez-Gazpio, I.; and Specia, L. 2017 · 2017
Cited alongside, same era.
First quora dataset release: Question pairs
Iyer, S.; Dandekar, N.; and Csernai, K. 2017 · 2017
Cited alongside, same era.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
Williams, A.; Nangia, N.; and Bowman, S. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Cited alongside, same era.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. R. 2019 · 2019
Cited alongside, same era.
PAWS: Paraphrase Adversaries from Word Scrambling
On the Ability and Limitations of Transformers to Recognize Formal Languages
Bhattamishra, S.; Ahuja, K.; and Goyal, N. 2020 · 2020
Later among the works it cites.
On the Importance of Pre-training Data Volume for Compact Language Models
Micheli, V.; d’Hoffschmidt, M.; and Fleuret, F. 2020 · 2020
Later among the works it cites.
Learning Music Helps You Read: Using Transfer to Study Linguistic Structure in Language Models
Papadimitriou, I.; and Jurafsky, D. 2020a · 2020
Later among the works it cites.
Learning Music Helps You Read: Using Transfer to Study Linguistic Structure in Language Models
Papadimitriou, I.; and Jurafsky, D. 2020b · 2020
Later among the works it cites.
A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource Languages
Suárez, P. J. O.; Romary, L.; and Sagot, B. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhang, Y.; Baldridge, J.; and He, L. 2019 · 2019
Cited alongside, same era.
Sinha, K.; Jia, R.; Hupkes, D.; Pineau, J.; Williams, A.; and Kiela, D. 2021 · 2021
Closest in time.