Well-read students learn better: On the importance of pre-training compact models
Original
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 1908
Earlier work this paper cites.
Teoria statistica delle classi ecalcolo delle probabilità
C.E. Bonferroni · 1936
Earlier work this paper cites.
Partialling out the spatial component of ecological variation
Daniel Borcard, Pierre Legendre, and Pierre Drapeau · 1992
Earlier work this paper cites.
Scaling laws for neural language models
Original
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeff Wu, and Dario Amodei · 2001
Earlier work this paper cites.
URIEL and lang2vec: Representing languages as typological, geographical, and phylogenetic vectors
Patrick Littell, David R. Mortensen, Ke Lin, Katherine Kairis, Carlisle Turner, and Lori Levin · 2002
Earlier work this paper cites.
Linguistically naïve != language independent: Why NLP needs linguistic typology
Emily M Bender · 2009
Earlier work this paper cites.
Haitian Creole language data
CMU · 2010
Earlier work this paper cites.
On achieving and evaluating language-independence in NLP
Emily M Bender · 2011
Earlier work this paper cites.
Building large monolingual dictionaries at the Leipzig corpora collection: From 100 to 200 languages
Dirk Goldhahn, Thomas Eckart, and Uwe Quasthoff · 2012
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson · 2012
Earlier work this paper cites.
Parallel data, tools and interfaces in OPUS
Jörg Tiedemann · 2012
Earlier work this paper cites.
WALS Online (v2020.3)
Matthew S. Dryer and Martin Haspelmath (eds.) · 2013
Earlier work this paper cites.
NCHLT isiXhosa Named Entity Annotated Corpus, 2016
Kholisa Podile and Roald Eiselen · 2016
Earlier work this paper cites.
The technology of web-texts collection of Russian minor languages
Lyudmila Zaydelman, Irina Krylova, and Boris Orekhov · 2016
Earlier work this paper cites.
On the relation between linguistic typology and (limitations of) multilingual language modeling
Daniela Gerz, Ivan Vulić, Edoardo Maria Ponti, Roi Reichart, and Anna Korhonen · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional Transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Asynchronous pipeline for processing huge corpora on medium to low resource infrastructures
Pedro Javier Ortiz Suárez, Benoît Sagot, and Laurent Romary · 2019
Earlier work this paper cites.
How multilingual is multilingual BERT?
Telmo Pires, Eva Schlinger, and Dan Garrette · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Energy and policy considerations for deep learning in NLP
Emma Strubell, Ananya Ganesh, and Andrew McCallum · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov · 2020
Earlier work this paper cites.
Emerging cross-lingual structure in pretrained language models
Alexis Conneau, Shijie Wu, Haoran Li, Luke Zettlemoyer, and Veselin Stoyanov · 2020
Earlier work this paper cites.
spaCy: Industrial-strength natural language processing in python
Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd · 2020
Earlier work this paper cites.
The Nunavut Hansard Inuktitut–English parallel corpus 3.0 with preliminary machine translation results, 2020
Eric Joanis, Rebecca Knowles, Roland Kuhn, Samuel Larkin, Patrick Littell, Chi-kiu Lo, Darlene Stewart, and Jeffrey Micher · 2020
Earlier work this paper cites.
The state and fate of linguistic diversity and inclusion in the NLP world
Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury · 2020
Earlier work this paper cites.