Fetching the paper…
Reading the bibliography…
How should text dataset sizes be compared across languages? Even for content-matched (parallel) corpora, UTF-8 encoded text can require a dramatically different number of bytes for different languages.
Fundamentals of grammatology
Peter T Daniels. 1990 · 1990
Earlier work this paper cites.
Segmental inventory size, word length, and communicative efficiency
Daniel Nettle. 1995 · 1995
Earlier work this paper cites.
The nature of the mental representation of radicals in Chinese: A priming study
Guosheng Ding, Danling Peng, and Marcus Taft. 2004 · 2004
Earlier work this paper cites.
AIC model selection using Akaike weights
Eric-Jan Wagenmakers and Simon Farrell. 2004 · 2004
Earlier work this paper cites.
Orthography to phonology and meaning: Comparisons across and within writing systems
Charles A Perfetti and Ying Liu. 2005 · 2005
Earlier work this paper cites.
Language identification of short text segments with n-gram models
Tommi Vatanen, Jaakko J. Väyrynen, and Sami Virpioja. 2010 · 2010
Earlier work this paper cites.
Chinese character decoding: a semantic bias?
Clay Williams and Thomas Bever. 2010 · 2010
Earlier work this paper cites.
Unicode over 60 percent of the web
Mark Davis. 2012 · 2012
Earlier work this paper cites.
Introducing language typology
Edith A Moravcsik. 2012 · 2012
Earlier work this paper cites.
PHOIBLE online
Steven Moran, Daniel McCloy, and Richard Wright. 2014 · 2014
Cited alongside, same era.
Byte-based neural machine translation
Marta R. Costa-jussà, Carlos Escolano, and José A. R. Fonollosa. 2017 · 2017
Cited alongside, same era.
XNLI: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Cited alongside, same era.
Wiki-40B: Multilingual language model dataset
Mandy Guo, Zihang Dai, Denny Vrandečić, and Rami Al-Rfou. 2020 · 2020
Cited alongside, same era.
XCOPA: A multilingual dataset for causal commonsense reasoning
Edoardo Maria Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulić, and Anna Korhonen. 2020 · 2020
BLOOM: A 176b-parameter open-access multilingual language model
Teven Le Scao, Angela Fan, Christopher Akiki, Elizabeth-Jane Pavlick, Suzana Ili’c, Daniel Hesslow, Roman Castagn’e, Alexandra Sasha Luccioni, Franccois Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Rose Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, Albert Villanova del Moral, Olatunji Ruwase, et al. 2022 · 2022
Later among the works it cites.
The Unicode Standard
Unicode Consortium. 2022 · 2022
Later among the works it cites.
Writing system and speaker metadata for 2,800+ language varieties
Daan van Esch, Tamar Lucassen, Sebastian Ruder, Isaac Caswell, and Clara Rivera. 2022 · 2022
Later among the works it cites.
ByT5: Towards a token-free future with pre-trained byte-to-byte models
Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel. 2022 · 2022
Later among the works it cites.
Byte-based multilingual NMT for endangered languages
Mengjiao Zhang and Jia Xu. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
No language left behind: Scaling human-centered machine translation
Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang. 2022 · 2022
Cited alongside, same era.
Few-shot learning with multilingual generative language models
Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O’Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona Diab, Veselin Stoyanov, and Xian Li. 2022 · 2022
Cited alongside, same era.
Do all languages cost the same? Tokenization in the era of commercial language models
Orevaoghene Ahia, Sachin Kumar, Hila Gonen, Jungo Kasai, David Mortensen, Noah Smith, and Yulia Tsvetkov. 2023 · 2023
Later among the works it cites.
Crosslingual generalization through multitask finetuning
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, and Colin Raffel. 2023 · 2023
Later among the works it cites.
Language model tokenizers introduce unfairness between languages
Aleksandar Petrov, Emanuele La Malfa, Philip Torr, and Adel Bibi. 2024 · 2024
Closest in time.