Fetching the paper…
Reading the bibliography…
Pre-trained text encoders such as BERT and its variants have recently achieved state-of-the-art performances on many NLP tasks.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, M. Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b · 1907
Earlier work this paper cites.
Multitask learning
Rich Caruana. 1997 · 1997
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B. Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
Mc-bert: Efficient language pre-training via a meta controller
Zhenhui Xu, Linyuan Gong, Guolin Ke, Di He, Shu xin Zheng, Liwei Wang, Jiang Bian, and T. Liu. 2020 · 2006
Earlier work this paper cites.
The third pascal recognizing textual entailment challenge
Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and William B. Dolan. 2007 · 2007
Earlier work this paper cites.
Clueweb09 data set
Jamie Callan, Mark Hoy, Changkuk Yoo, and Le Zhao. 2009 · 2009
Earlier work this paper cites.
On the transformer growth for progressive bert training
X. Gu, Liyuan Liu, H. Yu, Jing Li, Chen Chen, and J. Han. 2020 · 2010
Earlier work this paper cites.
The winograd schema challenge
H. Levesque, E. Davis, and L. Morgenstern. 2011 · 2011
Earlier work this paper cites.
English gigaword fifth edition ldc2011t07 (tech. rep.)
R Parker, D Graff, J Kong, K Chen, and K Maeda. 2011 · 2011
Earlier work this paper cites.
Progressively stacking 2.0: A multi-stage layerwise training method for bert training speedup
Cheng Yang, Shengnan Wang, Chao Yang, Yuechuan Li, Ru He, and Jingqiao Zhang. 2020 · 2011
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Y. Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Y. Zhu, Ryan Kiros, R. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba, and S. Fidler. 2015 · 2015
Cited alongside, same era.
Squad: 100, 000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Cited alongside, same era.
Semeval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Daniel M. Cer, Mona T. Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Cited alongside, same era.
First Quora dataset release: Question pairs
Shankar Iyer, Nikhil Dandekar, and Kornél Csernai. 2017 · 2017
Cited alongside, same era.
An overview of multi-task learning in deep neural networks
Sebastian Ruder. 2017 · 2017
Cited alongside, same era.
Common crawl
Common Crawl. 2019 · 2019
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Efficient training of bert by progressively stacking
Linyuan Gong, D. He, Zhuohan Li, T. Qin, Liwei Wang, and T. Liu. 2019 · 2019
Later among the works it cites.
Spanbert: Improving pre-training by representing and predicting spans
Mandar Joshi, Danqi Chen, Y. Liu, Daniel S. Weld, Luke Zettlemoyer, and Omer Levy. 2019 · 2019
Later among the works it cites.
Text summarization with pretrained encoders
Yang Liu and Mirella Lapata. 2019 · 2019
Later among the works it cites.
Mining entity synonyms with efficient neural set generation
J. Shen, Ruiliang Lyu, Xiang Ren, M. Vanni, Brian M. Sadler, and Jiawei Han. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, L. Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Cited alongside, same era.
Multi-task learning for email search ranking with auxiliary query clustering
J. Shen, Maryam Karimzadehgan, Michael Bendersky, Zhen Qin, and Donald Metzler. 2018 · 2018
Cited alongside, same era.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2018 · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018 · 2018
Cited alongside, same era.
Electra: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020a
Cited in the paper.
Xlnet: Generalized autoregressive pretraining for language understanding
Z. Yang, Zihang Dai, Yiming Yang, J. Carbonell, R. Salakhutdinov, and Quoc V. Le. 2019 · 2019
Later among the works it cites.
On losses for modern language models
Stephane T Aroca-Ouellette and F. Rudzicz. 2020 · 2020
Later among the works it cites.
Convbert: Improving bert with span-based dynamic convolution
Zihang Jiang, Weihao Yu, Daquan Zhou, Y. Chen, Jiashi Feng, and S. Yan. 2020 · 2020
Later among the works it cites.
Ernie 2.0: A continual pre-training framework for language understanding
Y. Sun, Shuohuan Wang, Yukun Li, Shikun Feng, Hao Tian, H. Wu, and Haifeng Wang. 2020 · 2020
Later among the works it cites.
Structbert: Incorporating language structures into pre-training for deep language understanding
Wei Wang, B. Bi, Ming Yan, Chen Wu, Zuyi Bao, Liwei Peng, and L. Si. 2020 · 2020
Later among the works it cites.