Fetching the paper…
Reading the bibliography…
In this paper, we study how the intrinsic nature of pre-training data contributes to the fine-tuned downstream performance.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2019 · 1910
Earlier work this paper cites.
Pre-training of deep bidirectional protein sequence representations with structural information
Seonwoo Min, Seunghyun Park, Siwon Kim, Hyun-Soo Choi, and Sungroh Yoon. 2019 · 1912
Earlier work this paper cites.
When do you need billions of words of pretraining data?
Yian Zhang, Alex Warstadt, Haau-Sing Li, and Samuel R Bowman. 2020 · 2011
Earlier work this paper cites.
Natural language models for predicting programming comments
Dana Movshovitz-Attias and William Cohen. 2013 · 2013
Earlier work this paper cites.
What do you learn from context? probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R Bowman, Dipanjan Das, et al. 2018 · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D Manning. 2019 · 2019
Cited alongside, same era.
Revealing the dark secrets of bert
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019 · 2019
Cited alongside, same era.
Open sesame: Getting inside bert’s linguistic knowledge
Yongjie Lin, Yi Chern Tan, and Robert Frank. 2019 · 2019
Cited alongside, same era.
Don’t stop pretraining: Adapt language models to domains and tasks
Bert rediscovers the classical nlp pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Later among the works it cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Later among the works it cites.
On the importance of pre-training data volume for compact language models
Vincent Micheli, Martin d’Hoffschmidt, and François Fleuret. 2020 · 2020
Closest in time.
Learning music helps you read: Using transfer to study linguistic structure in language models
Isabel Papadimitriou and Dan Jurafsky. 2020 · 2020
Closest in time.
A monolingual approach to contextualized word embeddings for mid-resource languages
Pedro Javier Ortiz Suárez, Laurent Romary, and Benoît Sagot. 2020 · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith
Cited in the paper.