Fetching the paper…
Reading the bibliography…
Large generative language models have been very successful for English, but other languages lag behind, in part due to data and computational limitations.
HuggingFace’s Transformers: State-of-the-art Natural Language Processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew. 2020 · 1910
Earlier work this paper cites.
Wietse de Vries, Andreas van Cranenburgh, Arianna Bisazza, Tommaso Caselli, Gertjan van Noord, and Malvina Nissim. 2019 · 1912
Earlier work this paper cites.
RobBERT: a Dutch RoBERTa-based Language Model
Pieter Delobelle, Thomas Winters, and Bettina Berendt. 2020 · 2001
Earlier work this paper cites.
What the [MASK]? Making Sense of Language-Specific BERT Models
Debora Nozza, Federico Bianchi, and Dirk Hovy. 2020 · 2003
Earlier work this paper cites.
GePpeTto Carves Italian into a Language Model
Lorenzo De Mattei, Michele Cafagna, Felice Dell’Orletta, Malvina Nissim, and Marco Guerini. 2020 · 2004
Earlier work this paper cites.
Are All Languages Created Equal in Multilingual BERT?
Shijie Wu and Mark Dredze. 2020 · 2005
Earlier work this paper cites.
TwNC: a Multifaceted Dutch News Corpus
Ordelman, Roeland J.F., de Jong, Franciska M.G., van Hessen, Adrianus J., and Hondorp, G.H.W. 2007 · 2007
Earlier work this paper cites.
The WaCky wide web: a collection of very large linguistically processed web-crawled corpora
Marco Baroni, Silvia Bernardini, Adriano Ferraresi, and Eros Zanchetta. 2009 · 2009
Earlier work this paper cites.
Exploiting Similarities among Languages for Machine Translation
Tomas Mikolov, Quoc V. Le, and Ilya Sutskever. 2013 · 2013
Earlier work this paper cites.
The Construction of a 500-Million-Word Reference Corpus of Contemporary Written Dutch
Nelleke Oostdijk, Martin Reynaert, Véronique Hoste, and Ineke Schuurman. 2013 · 2013
Earlier work this paper cites.
Improving Distributional Similarity with Lessons Learned from Word Embeddings
Omer Levy, Yoav Goldberg, and Ido Dagan. 2015 · 2015
Earlier work this paper cites.
Normalized Word Embedding and Orthogonal Transform for Bilingual Word Translation
Chao Xing, Dong Wang, Chao Liu, and Yiye Lin. 2015 · 2015
Earlier work this paper cites.
Learning principled bilingual mappings of word embeddings while preserving monolingual invariance
Mikel Artetxe, Gorka Labaka, and Eneko Agirre. 2016 · 2016
Earlier work this paper cites.
Transfer Learning for Low-Resource Neural Machine Translation
Barret Zoph, Deniz Yuret, Jonathan May, and Kevin Knight. 2016 · 2016
Cited alongside, same era.
Bag of Tricks for Efficient Text Classification
Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. 2017 · 2017
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba. 2017 · 2017
Cited alongside, same era.
Transfer Learning across Low-Resource, Related Languages for Neural Machine Translation
Toan Q. Nguyen and David Chiang. 2017 · 2017
Cited alongside, same era.
Cyclical Learning Rates for Training Neural Networks
Leslie N. Smith. 2017 · 2017
Cited alongside, same era.
How to (Properly) Evaluate Cross-Lingual Word Embeddings: On Strong Baselines, Comparative Analyses, and Some Misconceptions
Goran Glavaš, Robert Litschko, Sebastian Ruder, and Ivan Vulić. 2019 · 2019
Later among the works it cites.
Effective Cross-lingual Transfer of Neural Machine Translation Models without Shared Vocabularies
Yunsu Kim, Yingbo Gao, and Hermann Ney. 2019 · 2019
Later among the works it cites.
Choosing Transfer Languages for Cross-Lingual Learning
Yu-Hsiang Lin, Chian-Yu Chen, Jean Lee, Zirui Li, Yuyan Zhang, Mengzhou Xia, Shruti Rijhwani, Junxian He, Zhisong Zhang, Xuezhe Ma, Antonios Anastasopoulos, Patrick Littell, and Graham Neubig. 2019 · 2019
Later among the works it cites.
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 2019
Later among the works it cites.
PsychoPy2: Experiments in behavior made easy
Jonathan Peirce, Jeremy R. Gray, Sol Simpson, Michael MacAskill, Richard Höchenberger, Hiroyuki Sogo, Erik Kastman, and Jonas Kristoffer Lindeløv. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mikel Artetxe, Gorka Labaka, Eneko Agirre, and Kyunghyun Cho. 2018 · 2018
Cited alongside, same era.
Word Translation Without Parallel Data
Alexis Conneau, Guillaume Lample, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou. 2018 · 2018
Cited alongside, same era.
Loss in Translation: Learning Bilingual Word Mapping with a Retrieval Criterion
Armand Joulin, Piotr Bojanowski, Tomas Mikolov, Hervé Jégou, and Edouard Grave. 2018 · 2018
Cited alongside, same era.
Trivial Transfer Learning for Low-Resource Neural Machine Translation
Tom Kocmi and Ondřej Bojar. 2018 · 2018
Cited alongside, same era.
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu. 2018 · 2018
Cited alongside, same era.
RankME: Reliable Human Ratings for Natural Language Generation
Jekaterina Novikova, Ondřej Dušek, and Verena Rieser. 2018 · 2018
Cited alongside, same era.
On the Limitations of Unsupervised Bilingual Dictionary Induction
Anders Søgaard, Sebastian Ruder, and Ivan Vulić. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
A Survey of Cross-lingual Word Embedding Models
Sebastian Ruder, Ivan Vulić, and Anders Søgaard. 2019 · 2019
Later among the works it cites.
Energy and Policy Considerations for Deep Learning in NLP
Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2019 · 2019
Later among the works it cites.
On the Cross-lingual Transferability of Monolingual Representations
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020 · 2020
Closest in time.
Unsupervised Cross-lingual Representation Learning at Scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Closest in time.
Ethnologue: Languages of the World. Twenty-third edition
David M. Eberhard, Gary F. Simons, and Charles D. Fennig. 2020 · 2020
Closest in time.
LNMap: Departures from Isomorphic Assumption in Bilingual Lexicon Induction Through Non-Linear Mapping in Latent Space
Tasnim Mohiuddin, M Saiful Bari, and Shafiq Joty. 2020 · 2020
Closest in time.
What’s so special about BERT’s layers? A closer look at the NLP pipeline in monolingual and multilingual models
Wietse de Vries, Andreas van Cranenburgh, and Malvina Nissim. 2020 · 2020
Closest in time.