Fetching the paper…
Reading the bibliography…
Masked language models (MLMs) conventionally mask 15% of tokens due to the belief that more masking would leave insufficient context to learn good representations; this masking rate has been widely used, regardless of model sizes or masking strategies.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019 · 1909
Earlier work this paper cites.
The PASCAL recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini. 2005 · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B. Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
The second PASCAL recognising textual entailment challenge
Roy Bar Haim, Ido Dagan, Bill Dolan, Lisa Ferro, Danilo Giampiccolo, Bernardo Magnini, and Idan Szpektor. 2006 · 2006
Earlier work this paper cites.
The third PASCAL recognizing textual entailment challenge
Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and Bill Dolan. 2007 · 2007
Earlier work this paper cites.
The fifth PASCAL recognizing textual entailment challenge
Luisa Bentivogli, Peter Clark, Ido Dagan, and Danilo Giampiccolo. 2009 · 2009
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson. 2013 · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Gregory S. Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Earlier work this paper cites.
Sentence encoders on stilts: Supplementary training on intermediate labeled-data tasks
Jason Phang, Thibault Févry, and Samuel R Bowman. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional Transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Cited alongside, same era.
MASS: Masked sequence to sequence pre-training for language generation
Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2019 · 2019
Cited alongside, same era.
The bottom-up evolution of representations in the transformer: A study with machine translation and language modeling objectives
Elena Voita, Rico Sennrich, and Ivan Titov. 2019 · 2019
Cited alongside, same era.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020 · 2020
Later among the works it cites.
Masked language model scoring
Julian Salazar, Davis Liang, Toan Q. Nguyen, and Katrin Kirchhoff. 2020 · 2020
Later among the works it cites.
Blimp: The benchmark of linguistic minimal pairs for english
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R. Bowman. 2020 · 2020
Later among the works it cites.
DeBERTa: Decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021 · 2021
Later among the works it cites.
How to train BERT with an academic budget
Peter Izsak, Moshe Berchansky, and Omer Levy. 2021 · 2021
Later among the works it cites.
Does pretraining for summarization require knowledge transfer?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2019 · 2019
Cited alongside, same era.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2019 · 2019
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020 · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 2020
Cited alongside, same era.
ELECTRA: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020 · 2020
Cited alongside, same era.
SpanBERT: Improving pre-training by representing and predicting spans
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S. Weld, Luke Zettlemoyer, and Omer Levy. 2020 · 2020
Cited alongside, same era.
ALBERT: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020 · 2020
Cited alongside, same era.
Kundan Krishna, Jeffrey Bigham, and Zachary C. Lipton. 2021 · 2021
Later among the works it cites.
PMI-Masking: Principled masking of correlated spans
Yoav Levine, Barak Lenz, Opher Lieber, Omri Abend, Kevin Leyton-Brown, Moshe Tennenholtz, and Yoav Shoham. 2021 · 2021
Later among the works it cites.
Masked language modeling and the distributional hypothesis: Order word matters pre-training for little
Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams, and Douwe Kiela. 2021 · 2021
Later among the works it cites.
Frustratingly simple pretraining alternatives to masked language modeling
Atsuki Yamaguchi, George Chrysostomou, Katerina Margatina, and Nikolaos Aletras. 2021 · 2021
Later among the works it cites.
On the influence of masking policies in intermediate pre-training
Qinyuan Ye, Belinda Z. Li, Sinong Wang, Benjamin Bolte, Hao Ma, Wen-tau Yih, Xiang Ren, and Madian Khabsa. 2021 · 2021
Later among the works it cites.
How does the pre-training objective affect what large language models learn about linguistic properties?
Ahmed Alajrami and Nikolaos Aletras. 2022 · 2022
Closest in time.
Masked autoencoders as spatiotemporal learners
Christoph Feichtenhofer, Haoqi Fan, Yanghao Li, and Kaiming He. 2022 · 2022
Closest in time.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. 2022 · 2022
Closest in time.
Mask more and mask later: Efficient pre-training of masked language models by disentangling the [MASK] token
Baohao Liao, David Thulke, Sanjika Hewavitharana, Hermann Ney, and Christof Monz. 2022 · 2022
Closest in time.
VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. 2022 · 2022
Closest in time.
Image BERT pre-training with online tokenizer
Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen, Cihang Xie, Alan Yuille, and Tao Kong. 2022 · 2022
Closest in time.