Fetching the paper…
Reading the bibliography…
The core of self-supervised learning for pre-training language models includes pre-training task design as well as appropriate data augmentation.
ERNIE: Enhanced Representation through Knowledge Integration
Sun, Y.; Wang, S.; Li, Y.; Feng, S.; Chen, X.; Zhang, H.; Tian, X.; Zhu, D.; Tian, H.; and Wu, H. 2019 · 1904
Earlier work this paper cites.
Unified Language Model Pre-training for Natural Language Understanding and Generation
Dong, L.; Yang, N.; Wang, W.; Wei, F.; Liu, X.; Wang, Y.; Gao, J.; Zhou, M.; and Hon, H. 2019 · 1905
Earlier work this paper cites.
XLNet: Generalized Autoregressive Pretraining for Language Understanding
Yang, Z.; Dai, Z.; Yang, Y.; Carbonell, J. G.; Salakhutdinov, R.; and Le, Q. V. 2019 · 1906
Earlier work this paper cites.
SpanBERT: Improving Pre-training by Representing and Predicting Spans
Joshi, M.; Chen, D.; Liu, Y.; Weld, D. S.; Zettlemoyer, L.; and Levy, O. 2019 · 1907
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
StructBERT: Incorporating Language Structures into Pre-training for Deep Language Understanding
Wang, W.; Bi, B.; Yan, M.; Wu, C.; Bao, Z.; Peng, L.; and Si, L. 2019 · 1908
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V.; Debut, L.; Chaumond, J.; and Wolf, T. 2019 · 1910
Earlier work this paper cites.
“Cloze Procedure”: A New Tool for Measuring Readability
Taylor, W. L. 1953 · 1953
Earlier work this paper cites.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
Dodge, J.; Ilharco, G.; Schwartz, R.; Farhadi, A.; Hajishirzi, H.; and Smith, N. 2020 · 2002
Earlier work this paper cites.
Electra: Pre-training text encoders as discriminators rather than generators
Clark, K.; Luong, M.-T.; Le, Q. V.; and Manning, C. D. 2020b · 2003
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, W. B.; and Brockett, C. 2005 · 2005
Earlier work this paper cites.
Deberta: Decoding-enhanced bert with disentangled attention
He, P.; Liu, X.; Gao, J.; and Chen, W. 2020 · 2006
Earlier work this paper cites.
Self-Knowledge Distillation: A Simple Way for Better Generalization
Kim, K.; Ji, B.; Yoon, D.; and Hwang, S. 2020 · 2006
Earlier work this paper cites.
On the stability of fine-tuning bert: Misconceptions, explanations, and strong baselines
Mosbach, M.; Andriushchenko, M.; and Klakow, D. 2020 · 2006
Cited alongside, same era.
MC-BERT: Efficient Language Pre-Training via a Meta Controller
Xu, Z.; Gong, L.; Ke, G.; He, D.; Zheng, S.; Wang, L.; Bian, J.; and Liu, T. 2020b · 2006
Cited alongside, same era.
Revisiting few-sample BERT fine-tuning
Zhang, T.; Wu, F.; Katiyar, A.; Weinberger, K. Q.; and Artzi, Y. 2020 · 2006
Cited alongside, same era.
The third pascal recognizing textual entailment challenge
Giampiccolo, D.; Magnini, B.; Dagan, I.; and Dolan, W. B. 2007 · 2007
Cited alongside, same era.
First quora dataset release: Question pairs
Iyer, S.; Dandekar, N.; and Csernai, K. 2017 · 2017
Later among the works it cites.
Mean Teachers Are Better Role Models: Weight-Averaged Consistency Targets Improve Semi-Supervised Deep Learning Results
Tarvainen, A.; and Valpola, H. 2017 · 2017
Later among the works it cites.
Attention is All you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L. u.; and Polosukhin, I. 2017 · 2017
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A.; Nangia, N.; and Bowman, S. R. 2017 · 2017
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wan, Y.; Yang, B.; Wong, D. F.; Zhou, Y.; Chao, L. S.; Zhang, H.; and Chen, B. 2020 · 2010
Cited alongside, same era.
Noise-Contrastive Estimation of Unnormalized Statistical Models, with Applications to Natural Image Statistics
Gutmann, M. U.; and Hyvärinen, A. 2012 · 2012
Cited alongside, same era.
CLEAR: Contrastive Learning for Sentence Representation
Wu, Z.; Wang, S.; Gu, J.; Khabsa, M.; Sun, F.; and Ma, H. 2020 · 2012
Cited alongside, same era.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R.; Perelygin, A.; Wu, J.; Chuang, J.; Manning, C. D.; Ng, A. Y.; and Potts, C. 2013 · 2013
Cited alongside, same era.
Distilling the Knowledge in a Neural Network
Hinton, G.; Vinyals, O.; and Dean, J. 2015 · 2015
Cited alongside, same era.
Self-paced curriculum learning
Jiang, L.; Meng, D.; Zhao, Q.; Shan, S.; and Hauptmann, A. 2015 · 2015
Cited alongside, same era.
Temporal Ensembling for Semi-Supervised Learning
Laine, S.; and Aila, T. 2016 · 2016
Cited alongside, same era.
Squad: 100,000+ questions for machine comprehension of text
Rajpurkar, P.; Zhang, J.; Lopyrev, K.; and Liang, P. 2016 · 2016
Cited alongside, same era.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. R. 2018 · 2018
Later among the works it cites.
Neural network acceptability judgments
Warstadt, A.; Singh, A.; and Bowman, S. R. 2019 · 2019
Later among the works it cites.
Pre-Training Transformers as Energy-Based Cloze Models
Clark, K.; Luong, M.-T.; Le, Q.; and Manning, C. D. 2020a · 2020
Later among the works it cites.
Transformers: State-of-the-Art Natural Language Processing
Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; Davison, J.; Shleifer, S.; von Platen, P.; Ma, C.; Jernite, Y.; Plu, J.; Xu, C.; Scao, T. L.; Gugger, S.; Drame, M.; Lhoest, Q.; and Rush, A. M. 2020 · 2020
Later among the works it cites.
Revisiting Knowledge Distillation via Label Smoothing Regularization
Yuan, L.; Tay, F. E.; Li, G.; Wang, T.; and Feng, J. 2020 · 2020
Later among the works it cites.
Coco-lm: Correcting and contrasting text sequences for language model pretraining
Meng, Y.; Xiong, C.; Bajaj, P.; Tiwary, S.; Bennett, P.; Han, J.; and Song, X. 2021 · 2021
Closest in time.
SCRIPT: Self-Critic Pretraining of Transformers
Nijkamp, E.; Pang, B.; Wu, Y. N.; and Xiong, C. 2021 · 2021
Closest in time.