Fetching the paper…
Reading the bibliography…
Training monolingual language models for low and mid-resource languages is made challenging by limited and often inadequate pretraining data.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Julian Salazar, Davis Liang, Toan Q Nguyen, and Katrin Kirchhoff. 2019 · 1910
Earlier work this paper cites.
Benjamin van der Burgh and Suzan Verberne. 2019 · 1910
Earlier work this paper cites.
Wietse de Vries, Andreas van Cranenburgh, Arianna Bisazza, Tommaso Caselli, Gertjan van Noord, and Malvina Nissim. 2019 · 1912
Earlier work this paper cites.
Frysk wurdboek. 1 : Frysk-Nederlânsk , volume 1029
J. W. Zantema, editor. 1984 · 1984
Earlier work this paper cites.
Model compression
Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil. 2006 · 2006
Earlier work this paper cites.
Frysk Hânwurdboek , volume 1029 of Fryske Akademy
P. Duijff, F.J. van der Kuip, R. de Haan, and H. Sijens, editors. 2008 · 2008
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015 · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. 2016 · 2016
Earlier work this paper cites.
Increasing return on annotation investment: The automatic construction of a Universal Dependency treebank for Dutch
Gosse Bouma and Gertjan van Noord. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
OpenSubtitles2018: Statistical rescoring of sentence alignments in large, noisy parallel corpora
Pierre Lison, Jörg Tiedemann, and Milen Kouylekov. 2018 · 2018
Cited alongside, same era.
Advances in pre-training distributed word representations
Tomas Mikolov, Edouard Grave, Piotr Bojanowski, Christian Puhrsch, and Armand Joulin. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
How multilingual is multilingual BERT?
Telmo Pires, Eva Schlinger, and Dan Garrette. 2019 · 2019
Cited alongside, same era.
Translation artifacts in cross-lingual transfer learning
A primer in BERTology: What we know about how BERT works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020 · 2020
Later among the works it cites.
On negative interference in multilingual models: Findings and a meta-learning treatment
Zirui Wang, Zachary C. Lipton, and Yulia Tsvetkov. 2020 · 2020
Later among the works it cites.
Towards continual learning for multilingual machine translation via vocabulary substitution
Xavier Garcia, Noah Constant, Ankur Parikh, and Orhan Firat. 2021 · 2021
Later among the works it cites.
One size does not fit all: Finding the optimal subword sizes for FastText models across languages
Vít Novotný, Eniafe Festus Ayetiran, Dalibor Bačovský, Dávid Lupták, Michal Štefánik, and Petr Sojka. 2021 · 2021
Later among the works it cites.
A statistical extension of byte-pair encoding
David Vilar and Marcello Federico. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mikel Artetxe, Gorka Labaka, and Eneko Agirre. 2020a · 2020
Cited alongside, same era.
German’s next language model
Branden Chan, Stefan Schweter, and Timo Möller. 2020 · 2020
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Cited alongside, same era.
RobBERT: a Dutch RoBERTa-based Language Model
Pieter Delobelle, Thomas Winters, and Bettina Berendt. 2020 · 2020
Cited alongside, same era.
CamemBERT: a tasty French language model
Louis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont, Laurent Romary, Éric de la Clergerie, Djamé Seddah, and Benoît Sagot. 2020 · 2020
Cited alongside, same era.
BPE-dropout: Simple and effective subword regularization
Ivan Provilkov, Dmitrii Emelianenko, and Elena Voita. 2020 · 2020
Cited alongside, same era.
On the cross-lingual transferability of monolingual representations
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020b
Cited in the paper.
Wietse de Vries, Martijn Bartelds, Malvina Nissim, and Martijn Wieling. 2021 · 2021
Later among the works it cites.
As good as new. how to successfully recycle English GPT-2 to make models for other languages
Wietse de Vries and Malvina Nissim. 2021 · 2021
Later among the works it cites.
SICK-NL: A dataset for Dutch natural language inference
Gijs Wijnholds and Michael Moortgat. 2021 · 2021
Later among the works it cites.
Taming large lexicons : translating clinical text using medical ontologies and sentence templates
François Remy, Peter De Jaeger, and Kris Demuynck. 2022 · 2022
Later among the works it cites.
Online language modelling training pipeline
Tristan Thrush and Muhtasham Muhtasham Oblokulov. 2022 · 2022
Later among the works it cites.
Revisiting machine translation for cross-lingual classification
Mikel Artetxe, Vedanuj Goswami, Shruti Bhosale, Angela Fan, and Luke Zettlemoyer. 2023 · 2023
Closest in time.