Fetching the paper…
Reading the bibliography…
Instead of pretraining multilingual language models from scratch, a more efficient method is to adapt existing pretrained language models (PLMs) to new languages via vocabulary extension and continued pretraining.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Quantifying the carbon emissions of machine learning
Alexandre Lacoste, Alexandra Luccioni, Victor Schmidt, and Thomas Dandres. 2019 · 1910
Earlier work this paper cites.
Cross-lingual name tagging and linking for 282 languages
Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji. 2017 · 1958
Earlier work this paper cites.
From english to foreign languages: Transferring pre-trained language models
Ke Tran. 2020 · 2002
Earlier work this paper cites.
Semantic maps and the typology of colexification
Alexandre François. 2008 · 2008
Earlier work this paper cites.
Creating a massively parallel Bible corpus
Thomas Mayer and Michael Cysouw. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell. 2017 · 2017
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Earlier work this paper cites.
Mixed precision training
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory F. Diamos, Erich Elsen, David García, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu. 2018 · 2018
Earlier work this paper cites.
Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond
Mikel Artetxe and Holger Schwenk. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Unicoder: A universal language encoder by pre-training with multiple cross-lingual tasks
Haoyang Huang, Yaobo Liang, Nan Duan, Ming Gong, Linjun Shou, Daxin Jiang, and Ming Zhou. 2019 · 2019
Earlier work this paper cites.
How multilingual is multilingual BERT?
Telmo Pires, Eva Schlinger, and Dan Garrette. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Earlier work this paper cites.
Improving pre-trained multilingual model with vocabulary expansion
Hai Wang, Dian Yu, Kai Sun, Jianshu Chen, and Dong Yu. 2019 · 2019
Earlier work this paper cites.
On the cross-lingual transferability of monolingual representations
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020 · 2020
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Earlier work this paper cites.
XTREME: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020 · 2020
Earlier work this paper cites.
SimAlign: High quality word alignments without parallel training data using static and contextualized embeddings
Masoud Jalili Sabet, Philipp Dufter, François Yvon, and Hinrich Schütze. 2020 · 2020
Earlier work this paper cites.
Cross-lingual ability of multilingual BERT: an empirical study
Karthikeyan K, Zihan Wang, Stephen Mayhew, and Dan Roth. 2020 · 2020
Earlier work this paper cites.
IndicNLPSuite: Monolingual corpora, evaluation benchmarks and pre-trained multilingual language models for Indian languages
Divyanshu Kakwani, Anoop Kunchukuttan, Satish Golla, Gokul N.C., Avik Bhattacharyya, Mitesh M. Khapra, and Pratyush Kumar. 2020 · 2020
Cited alongside, same era.
ALBERT: A lite BERT for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020 · 2020
Cited alongside, same era.
Multilingual denoising pre-training for neural machine translation
Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020 · 2020
Cited alongside, same era.
exBERT: Extending pre-trained models with domain-specific vocabulary under constrained training resources
Wen Tai, H. T. Kung, Xin Dong, Marcus Comiter, and Chang-Fu Kuo. 2020 · 2020
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
WECHSEL: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language models
Benjamin Minixhofer, Fabian Paischer, and Navid Rekabsaz. 2022 · 2022
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022 · 2022
Later among the works it cites.
mgpt: Few-shot learners go multilingual
Oleh Shliazhko, Alena Fenogenova, Maria Tikhonova, Vladislav Mikhailov, Anastasia Kozlova, and Tatiana Shavrina. 2022 · 2022
Later among the works it cites.
Scale efficiently: Insights from pretraining and finetuning transformers
Yi Tay, Mostafa Dehghani, Jinfeng Rao, William Fedus, Samira Abnar, Hyung Won Chung, Sharan Narang, Dani Yogatama, Ashish Vaswani, and Donald Metzler. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Rethinking embedding coupling in pre-trained language models
Hyung Won Chung, Thibault Févry, Henry Tsai, Melvin Johnson, and Sebastian Ruder. 2021 · 2021
Cited alongside, same era.
Universal Dependencies
Marie-Catherine de Marneffe, Christopher D. Manning, Joakim Nivre, and Daniel Zeman. 2021 · 2021
Cited alongside, same era.
As good as new. how to successfully recycle English GPT-2 to make models for other languages
Wietse de Vries and Malvina Nissim. 2021 · 2021
Cited alongside, same era.
Initializing new word embeddings for pretrained language models
John Hewitt. 2021 · 2021
Cited alongside, same era.
AVocaDo: Strategy for adapting vocabulary to downstream domain
Jimin Hong, TaeHee Kim, Hyesu Lim, and Jaegul Choo. 2021 · 2021
Cited alongside, same era.
Small data? no problem! exploring the viability of pretrained multilingual language models for low-resourced languages
Kelechi Ogueji, Yuxin Zhu, and Jimmy Lin. 2021 · 2021
Cited alongside, same era.
Multilingual BERT post-pretraining alignment
Lin Pan, Chung-Wei Hang, Haode Qi, Abhishek Shah, Saloni Potdar, and Mo Yu. 2021 · 2021
Cited alongside, same era.
Xinyi Wang, Sebastian Ruder, and Graham Neubig. 2022 · 2022
Later among the works it cites.
Taxonomy of risks posed by language models
Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, Courtney Biles, Sasha Brown, Zac Kenton, Will Hawkins, Tom Stepleton, Abeba Birhane, Lisa Anne Hendricks, Laura Rimell, William Isaac, Julia Haas, Sean Legassick, Geoffrey Irving, and Iason Gabriel. 2022 · 2022
Later among the works it cites.
Improving language plasticity via pretraining with active forgetting
Yihong Chen, Kelly Marchisio, Roberta Raileanu, David Ifeoluwa Adelani, Pontus Stenetor, Sebastian Riedel, and Mikel Artetx. 2023 · 2023
Closest in time.
FOCUS: Effective embedding initialization for monolingual specialization of multilingual models
Konstantin Dobler and Gerard de Melo. 2023 · 2023
Closest in time.
Continual pre-training of large language models: How to (re) warm your model?
Kshitij Gupta, Benjamin Thérien, Adam Ibrahim, Mats L Richter, Quentin Anthony, Eugene Belilovsky, Irina Rish, and Timothée Lesort. 2023 · 2023
Closest in time.
Glot500: Scaling multilingual corpora and language models to 500 languages
Ayyoob ImaniGooghari, Peiqin Lin, Amir Hossein Kargaran, Silvia Severini, Masoud Jalili Sabet, Nora Kassner, Chunlan Ma, Helmut Schmid, André Martins, François Yvon, and Hinrich Schütze. 2023 · 2023
Closest in time.
Isotropic representation can improve zero-shot cross-lingual transfer on multilingual language models
Yixin Ji, Jikai Wang, Juntao Li, Hai Ye, and Min Zhang. 2023 · 2023
Closest in time.
XLM-V: Overcoming the vocabulary bottleneck in multilingual masked language models
Davis Liang, Hila Gonen, Yuning Mao, Rui Hou, Naman Goyal, Marjan Ghazvininejad, Luke Zettlemoyer, and Madian Khabsa. 2023 · 2023
Closest in time.
Crosslingual transfer learning for low-resource languages based on multilingual colexification graphs
Yihong Liu, Haotian Ye, Leonie Weissweiler, Renhao Pei, and Hinrich Schuetze. 2023a · 2023
Closest in time.
Taxi1500: A multilingual dataset for text classification in 1500 languages
Chunlan Ma, Ayyoob ImaniGooghari, Haotian Ye, Ehsaneddin Asgari, and Hinrich Schütze. 2023 · 2023
Closest in time.
Mini-model adaptation: Efficiently extending pretrained models to new languages via aligned shallow training
Kelly Marchisio, Patrick Lewis, Yihong Chen, and Mikel Artetxe. 2023 · 2023
Closest in time.
Subspace chronicles: How linguistic information emerges, shifts and interacts during language model training
Max Müller-Eberstein, Rob van der Goot, Barbara Plank, and Ivan Titov. 2023 · 2023
Closest in time.
Entropy-guided vocabulary augmentation of multilingual language models for low-resource tasks
Arijit Nag, Bidisha Samanta, Animesh Mukherjee, Niloy Ganguly, and Soumen Chakrabarti. 2023 · 2023
Closest in time.
A study of conceptual language similarity: comparison and evaluation
Haotian Ye, Yihong Liu, and Hinrich Schütze. 2023 · 2023
Closest in time.
A comprehensive survey on pretrained foundation models: A history from bert to chatgpt
Ce Zhou, Qian Li, Chen Li, Jun Yu, Yixin Liu, Guangjing Wang, Kai Zhang, Cheng Ji, Qiben Yan, Lifang He, et al. 2023 · 2023
Closest in time.
Yihong Liu, Chunlan Ma, Haotian Ye, and Hinrich Schütze. 2024 · 2024
Closest in time.