Fetching the paper…
Reading the bibliography…
A major consideration in multilingual language modeling is how to best represent languages with diverse vocabularies and scripts.
Understanding Morphology
M. Haspelmath and A.D. Sims. 2010 · 2010
Earlier work this paper cites.
The Unicode Standard
The Unicode Consortium. 2011 · 2011
Earlier work this paper cites.
Morfessor 2.0: Toolkit for statistical morphological segmentation
Peter Smit, Sami Virpioja, Stig-Arne Grönroos, and Mikko Kurimo. 2014 · 2014
Earlier work this paper cites.
Neural Machine Translation of Rare Words with Subword Units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Are All Languages Equally Hard to Language-Model?
Ryan Cotterell, S. J. Mielke, Jason Eisner, and Brian Roark. 2018 · 2018
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Earlier work this paper cites.
Word translation without parallel data
Guillaume Lample, Alexis Conneau, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou. 2018 · 2018
Earlier work this paper cites.
Morphological and Language-agnostic Word Segmentation for NMT
Dominik Machácek, Jonás Vidra, and Ondrej Bojar. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Improving Multilingual Models with Language-clustered Vocabularies
Hyung Won Chung, Dan Garrette, Kiat Chuan Tan, and Jason Riesa. 2020 · 2020
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Earlier work this paper cites.
The state and fate of linguistic diversity and inclusion in the NLP world
Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020 · 2020
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Earlier work this paper cites.
Morphology Matters: A Multilingual Language Modeling Analysis
Hyunji Hayley Park, Katherine J. Zhang, Coleman Haley, Kenneth Steimel, Han Liu, and Lane Schwartz. 2021 · 2021
Earlier work this paper cites.
UNKs everywhere: Adapting multilingual language models to new scripts
Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych, and Sebastian Ruder. 2021 · 2021
Cited alongside, same era.
mT5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021 · 2021
Cited alongside, same era.
Allocating Large Vocabulary Capacity for Cross-lingual Language Model Pre-training
Bo Zheng, Li Dong, Shaohan Huang, Saksham Singhal, Wanxiang Che, Ting Liu, Xia Song, and Furu Wei. 2021 · 2021
Cited alongside, same era.
MasakhaNER 2.0: Africa-centric transfer learning for named entity recognition
David Adelani, Graham Neubig, Sebastian Ruder, Shruti Rijhwani, Michael Beukman, Chester Palen-Michel, Constantine Lignos, Jesujoba Alabi, Shamsuddeen Muhammad, Peter Nabende, Cheikh M. Bamba Dione, Andiswa Bukula, Rooweither Mabuya, Bonaventure F. P. Dossou, Blessing Sibanda, Happy Buzaaba, Jonathan Mukiibi, Godson Kalipe, Derguene Mbaye, Amelia Taylor, Fatoumata Kabore, Chris Chinenye Emezue, Anuoluwapo Aremu, Perez Ogayo, Catherine Gitau, Edwin Munkoh-Buabeng, Victoire Memdjokam Koagne, Allahsera Auguste Tapo, Tebogo Macucwa, Vukosi Marivate, Mboning Tchiaze Elvis, Tajuddeen Gwadabe, Tosin Adewumi, Orevaoghene Ahia, Joyce Nakatumba-Nabende, Neo Lerato Mokono, Ignatius Ezeani, Chiamaka Chukwuneke, Mofetoluwa Oluwaseun Adeyemi, Gilles Quentin Hacheme, Idris Abdulmumin, Odunayo Ogundepo, Oreen Yousuf, Tatiana Moteu, and Dietrich Klakow. 2022 · 2022
UniMax: Fairer and More Effective Language Sampling for Large-scale Multilingual Pretraining
Hyung Won Chung, Xavier Garcia, Adam Roberts, Yi Tay, Orhan Firat, Sharan Narang, and Noah Constant. 2023 · 2023
Later among the works it cites.
Effects of sub-word segmentation on performance of transformer language models
Jue Hou, Anisia Katinskaia, Anh-Duc Vu, and Roman Yangarber. 2023 · 2023
Later among the works it cites.
XLM-V: Overcoming the Vocabulary Bottleneck in Multilingual Masked Language Models
Davis Liang, Hila Gonen, Yuning Mao, Rui Hou, Naman Goyal, Marjan Ghazvininejad, Luke Zettlemoyer, and Madian Khabsa. 2023 · 2023
Later among the works it cites.
Tokenization Impacts Multilingual Language Modeling: Assessing Vocabulary Allocation and Overlap Across Languages
Tomasz Limisiewicz, Jirí Balhar, and David Marecek. 2023 · 2023
Later among the works it cites.
Efficient Transformers with Dynamic Token Pooling
Piotr Nawrot, Jan Chorowski, Adrian Lancucki, and Edoardo Maria Ponti. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Language Contamination Helps Explains the Cross-lingual Capabilities of English Pretrained Models
Terra Blevins and Luke Zettlemoyer. 2022 · 2022
Cited alongside, same era.
Canine: Pre-training an efficient tokenization-free encoder for language representation
Jonathan H. Clark, Dan Garrette, Iulia Turc, and John Wieting. 2022 · 2022
Cited alongside, same era.
A balanced data approach for evaluating cross-lingual transfer: Mapping the linguistic blood bank
Dan Malkin, Tomasz Limisiewicz, and Gabriel Stanovsky. 2022 · 2022
Cited alongside, same era.
Scaling Up Models and Data with $\texttt{t5x}$ and $\texttt{seqio}$
Adam Roberts, Hyung Won Chung, Anselm Levskaya, Gaurav Mishra, James Bradbury, Daniel Andor, Sharan Narang, Brian Lester, Colin Gaffney, Afroz Mohiuddin, Curtis Hawthorne, Aitor Lewkowycz, Alex Salcianu, Marc van Zee, Jacob Austin, Sebastian Goodman, Livio Baldini Soares, Haitang Hu, Sasha Tsvyashchenko, Aakanksha Chowdhery, Jasmijn Bastings, Jannis Bulian, Xavier Garcia, Jianmo Ni, Andrew Chen, Kathleen Kenealy, Jonathan H. Clark, Stephan Lee, Dan Garrette, James Lee-Thorp, Colin Raffel, Noam Shazeer, Marvin Ritter, Maarten Bosma, Alexandre Passos, Jeremy Maitin-Shepard, Noah Fiedel, Mark Omernick, Brennan Saeta, Ryan Sepassi, Alexander Spiridonov, Joshua Newlan, and Andrea Gesmundo. 2022 · 2022
Cited alongside, same era.
Charformer: Fast Character Transformers via Gradient-based Subword Tokenization
Yi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Prakash Gupta, Hyung Won Chung, Dara Bahri, Zhen Qin, Simon Baumgartner, Cong Yu, and Donald Metzler. 2022 · 2022
Cited alongside, same era.
No Language Left Behind: Scaling Human-centered Machine Translation
NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang. 2022 · 2022
Cited alongside, same era.
ByT5: Towards a Token-free Future with Pre-trained Byte-to-byte Models
Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel. 2022 · 2022
Cited alongside, same era.
Towards universal segmentations: UniSegments 1.0
Zdeněk Žabokrtský, Niyati Bafna, Jan Bodnár, Lukáš Kyjánek, Emil Svoboda, Magda Ševčíková, and Jonáš Vidra. 2022 · 2022
Cited alongside, same era.
Language Model Tokenizers Introduce Unfairness Between Languages
Aleksandar Petrov, Emanuele La Malfa, Philip H. S. Torr, and Adel Bibi. 2023 · 2023
Later among the works it cites.
XTREME-UP: A User-centric Scarce-data Benchmark for Under-represented Languages
Sebastian Ruder, Jonathan H. Clark, Alexander Gutkin, Mihir Kale, Min Ma, Massimo Nicosia, Shruti Rijhwani, Parker Riley, Jean Michel A. Sarr, Xinyi Wang, John Wieting, Nitish Gupta, Anna Katanova, Christo Kirov, Dana L. Dickinson, Brian Roark, Bidisha Samanta, Connie Tao, David Ifeoluwa Adelani, Vera Axelrod, Isaac Caswell, Colin Cherry, Dan Garrette, R. Reeve Ingle, Melvin Johnson, Dmitry Panteleev, and Partha Talukdar. 2023 · 2023
Later among the works it cites.
Language Modelling with Pixels
Phillip Rust, Jonas F. Lotz, Emanuele Bugliarello, Elizabeth Salesky, Miryam de Lhoneux, and Desmond Elliott. 2023 · 2023
Later among the works it cites.
A Multi-dimensional Evaluation of Tokenizer-free Multilingual Pretrained Models
Jimin Sun, Patrick Fernandes, Xinyi Wang, and Graham Neubig. 2023 · 2023
Later among the works it cites.
MEGABYTE: Predicting Million-byte Sequences with Multiscale Transformers
Lili Yu, Daniel Simig, Colin Flaherty, Armen Aghajanyan, Luke Zettlemoyer, and Mike Lewis. 2023 · 2023
Later among the works it cites.
A Formal Perspective on Byte-pair Encoding
Vilém Zouhar, Clara Meister, Juan Luis Gastaldi, Li Du, Tim Vieira, Mrinmaya Sachan, and Ryan Cotterell. 2023 · 2023
Later among the works it cites.
A bit of a problem: Measurement disparities in dataset sizes across languages
Catherine Arnett, Tyler A. Chang, and Benjamin K. Bergen. 2024 · 2024
Closest in time.
GlotScript: A resource and tool for low resource writing system identification
Amir Hossein Kargaran, François Yvon, and Hinrich Schütze. 2024 · 2024
Closest in time.