Fetching the paper…
Reading the bibliography…
For many low-resource languages, the only available language models are large multilingual models trained on many languages simultaneously.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020a · 1901
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020b · 1901
Earlier work this paper cites.
Pratik Joshi, Christain Barnes, Sebastin Santy, Simran Khanuja, Sanket Shah, Anirudh Srinivasan, Satwik Bhattamishra, Sunayana Sitaram, Monojit Choudhury, and Kalika Bali. 2019 · 1912
Earlier work this paper cites.
Teoria statistica delle classi ecalcolo delle probabilità
C.E. Bonferroni. 1936 · 1936
Earlier work this paper cites.
The makerere radio speech corpus: A Luganda radio corpus for automatic speech recognition
Jonathan Mukiibi, Andrew Katumba, Joyce Nakatumba-Nabende, Ali Hussein, and Joshua Meyer. 2022 · 1954
Earlier work this paper cites.
An algorithm for the machine calculation of complex Fourier series
James W. Cooley and John W. Tukey. 1965 · 1965
Earlier work this paper cites.
The Theory of Parsing, Translation and Compiling , volume 1
Alfred V. Aho and Jeffrey D. Ullman. 1972 · 1972
Earlier work this paper cites.
Alternation
Ashok K. Chandra, Dexter C. Kozen, and Larry J. Stockmeyer. 1981 · 1981
Earlier work this paper cites.
Publications Manual
American Psychological Association. 1983 · 1983
Earlier work this paper cites.
Estimation of probabilities from sparse data for the language model component of a speech recognizer
Slava Katz. 1987 · 1987
Earlier work this paper cites.
Partialling out the spatial component of ecological variation
Daniel Borcard, Pierre Legendre, and Pierre Drapeau. 1992 · 1992
Earlier work this paper cites.
Algorithms on Strings, Trees and Sequences
Dan Gusfield. 1997 · 1997
Earlier work this paper cites.
A framework for learning predictive structures from multiple tasks and unlabeled data
Rie Kubota Ando and Tong Zhang. 2005 · 2005
Earlier work this paper cites.
Establishing baselines for text classification in low-resource languages
Jan Christian Blaise Cruz and Charibeth Cheng. 2020 · 2005
Earlier work this paper cites.
Unsupervised cross-lingual representation learning for speech recognition
Alexis Conneau, Alexei Baevski, Ronan Collobert, Abdelrahman Mohamed, and Michael Auli. 2020a · 2006
Earlier work this paper cites.
Low-resource languages: A review of past work and future challenges
Alexandre Magueresse, Vincent Carles, and Evan Heetderks. 2020 · 2006
Earlier work this paper cites.
Scalable training of L 1 L_{1} -regularized log-linear models
Galen Andrew and Jianfeng Gao. 2007 · 2007
Earlier work this paper cites.
Large language models in machine translation
Thorsten Brants, Ashok C. Popat, Peng Xu, Franz J. Och, and Jeffrey Dean. 2007 · 2007
Earlier work this paper cites.
Linguistically naïve != language independent: Why NLP needs linguistic typology
Emily M Bender. 2009 · 2009
Earlier work this paper cites.
Haitian Creole language data
CMU. 2010 · 2010
Earlier work this paper cites.
On achieving and evaluating language-independence in NLP
Emily M Bender. 2011 · 2011
Earlier work this paper cites.
W2C – web to corpus – corpora
Martin Majliš. 2011 · 2011
Earlier work this paper cites.
The Kyoto free translation task
Graham Neubig. 2011 · 2011
Earlier work this paper cites.
Building large monolingual dictionaries at the Leipzig corpora collection: From 100 to 200 languages
Dirk Goldhahn, Thomas Eckart, and Uwe Quasthoff. 2012 · 2012
Earlier work this paper cites.
How good are typological distances for determining genealogical relationships among languages?
Taraka Rama and Prasanth Kolachina. 2012 · 2012
Earlier work this paper cites.
Parallel data, tools and interfaces in OPUS
Jörg Tiedemann. 2012 · 2012
Earlier work this paper cites.
WALS Online (v2020.3)
Matthew S. Dryer and Martin Haspelmath, editors. 2013 · 2013
Earlier work this paper cites.
Creating a massively parallel Bible corpus
Thomas Mayer and Michael Cysouw. 2014 · 2014
Earlier work this paper cites.
Yara parser: A fast and accurate dependency parser
Mohammad Sadegh Rasooli and Joel R. Tetreault. 2015 · 2015
Earlier work this paper cites.
A large-scale multilingual disambiguation of glosses
José Camacho-Collados, Claudio Delli Bovi, Alessandro Raganato, and Roberto Navigli. 2016 · 2016
Earlier work this paper cites.
Selection criteria for low resource language programs
Christopher Cieri, Mike Maxwell, Stephanie Strassel, and Jennifer Tracey. 2016 · 2016
Earlier work this paper cites.
NCHLT isiXhosa Named Entity Annotated Corpus
Kholisa Podile and Roald Eiselen. 2016 · 2016
Earlier work this paper cites.
The technology of web-texts collection of Russian minor languages
Lyudmila Zaydelman, Irina Krylova, and Boris Orekhov. 2016 · 2016
Earlier work this paper cites.
URIEL and lang2vec: Representing languages as typological, geographical, and phylogenetic vectors
Patrick Littell, David R. Mortensen, Ke Lin, Katherine Kairis, Carlisle Turner, and Lori Levin. 2017 · 2017
Earlier work this paper cites.
Tilde MODEL - multilingual open data for EU languages
Roberts Rozis and Raivis Skadiņš. 2017 · 2017
Earlier work this paper cites.
Shami: A corpus of Levantine Arabic dialects
Kathrein Abu Kwaik, Motaz Saad, Stergios Chatzikyriakidis, and Simon Dobnik. 2018 · 2018
Earlier work this paper cites.
Developing new linguistic resources and tools for the Galician language
Rodrigo Agerri, Xavier Gómez Guinovart, German Rigau, and Miguel Anxo Solla Portela. 2018 · 2018
Earlier work this paper cites.
DART: A large dataset of dialectal Arabic tweets
Israa Alsarsour, Esraa Mohamed, Reem Suwaileh, and Tamer Elsayed. 2018 · 2018
Earlier work this paper cites.
Arabic dialect identification in the context of bivalency and code-switching
Mahmoud El-Haj, Paul Rayson, and Mariam Aboelezz. 2018 · 2018
Earlier work this paper cites.
On the relation between linguistic typology and (limitations of) multilingual language modeling
Daniela Gerz, Ivan Vulić, Edoardo Maria Ponti, Roi Reichart, and Anna Korhonen. 2018 · 2018
Earlier work this paper cites.
Untranslatability Goes Global
Suzanne Jill Levine and Katie Lateef-Jan. 2018 · 2018
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Earlier work this paper cites.
The IIT Bombay English-Hindi parallel corpus
Anoop Kunchukuttan, Pratik Mehta, and Pushpak Bhattacharyya. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Earlier work this paper cites.
Parallel corpora for bi-directional statistical machine translation for seven Ethiopian language pairs
Solomon Teferra Abate, Michael Melese, Martha Yifiru Tachbelie, Million Meshesha, Solomon Atinafu, Wondwossen Mulugeta, Yaregal Assabie, Hafte Abera, Binyam Ephrem, Tewodros Abebe, Wondimagegnhue Tsegaye, Amanuel Lemma, Tsegaye Andargie, and Seifedin Shifaw. 2018 · 2018
Earlier work this paper cites.
The# benderrule: On naming the languages we study and why it matters
Emily Bender. 2019 · 2019
Earlier work this paper cites.
Cross-lingual language model pretraining
Alexis Conneau and Guillaume Lample. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional Transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
hULMonA: The universal language model in Arabic
Obeida ElJundi, Wissam Antoun, Nour El Droubi, Hazem Hajj, Wassim El-Hajj, and Khaled Shaban. 2019 · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Earlier work this paper cites.
Asynchronous pipeline for processing huge corpora on medium to low resource infrastructures
Pedro Javier Ortiz Suárez, Benoît Sagot, and Laurent Romary. 2019 · 2019
Earlier work this paper cites.
How multilingual is multilingual BERT?
Telmo Pires, Eva Schlinger, and Dan Garrette. 2019 · 2019
Earlier work this paper cites.
Alberto: Modeling italian social media language with bert
Marco Polignano, Valerio Basile, Pierpaolo Basile, Marco de Gemmis, and Giovanni Semeraro. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
Energy and policy considerations for deep learning in NLP
Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2019 · 2019
Earlier work this paper cites.
Well-read students learn better: On the importance of pre-training compact models
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
TICO-19: the translation initiative for COvid-19
Antonios Anastasopoulos, Alessandro Cattelan, Zi-Yi Dou, Marcello Federico, Christian Federmann, Dmitriy Genzel, Franscisco Guzmán, Junjie Hu, Macduff Hughes, Philipp Koehn, Rosie Lazar, Will Lewis, Graham Neubig, Mengmeng Niu, Alp Öktem, Eric Paquin, Grace Tang, and Sylwia Tur. 2020 · 2020
Earlier work this paper cites.
On the cross-lingual transferability of monolingual representations
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020 · 2020
Earlier work this paper cites.
ParaCrawl: Web-scale acquisition of parallel corpora
Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz Rojas, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Elsa Sarrías, Marek Strelec, Brian Thompson, William Waites, Dion Wiggins, and Jaume Zaragoza. 2020 · 2020
Earlier work this paper cites.
Measuring diachronic language distance using perplexity: Application to english, portuguese, and spanish
José Ramom Pichel Campos, Pablo Gamallo Otero, and Iñaki Alegria Loinaz. 2020 · 2020
Earlier work this paper cites.
Language ID in the wild: Unexpected challenges on the path to a thousand-language web text corpus
Isaac Caswell, Theresa Breiner, Daan van Esch, and Ankur Bapna. 2020 · 2020
Earlier work this paper cites.
Finding universal grammatical relations in multilingual bert
Ethan A Chi, John Hewitt, and Christopher D Manning. 2020 · 2020
Earlier work this paper cites.
Mapping languages: the corpus of global language use
Jonathan Dunn. 2020 · 2020
Earlier work this paper cites.
Habibi - a multi dialect multi national Arabic song lyrics corpus
Mahmoud El-Haj. 2020 · 2020
Earlier work this paper cites.
spaCy: Industrial-strength natural language processing in python
Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd. 2020 · 2020
Earlier work this paper cites.
The Nunavut Hansard Inuktitut–English parallel corpus 3.0 with preliminary machine translation results
Eric Joanis, Rebecca Knowles, Roland Kuhn, Samuel Larkin, Patrick Littell, Chi-kiu Lo, Darlene Stewart, and Jeffrey Micher. 2020 · 2020
Earlier work this paper cites.
IndicNLPSuite: Monolingual corpora, evaluation benchmarks and pre-trained multilingual language models for Indian languages
Divyanshu Kakwani, Anoop Kunchukuttan, Satish Golla, Gokul N.C., Avik Bhattacharyya, Mitesh M. Khapra, and Pratyush Kumar. 2020 · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeff Wu, and Dario Amodei. 2020 · 2020
Earlier work this paper cites.
Cross-lingual ability of multilingual BERT: An empirical study
Karthikeyan, Zihan Wang, Stephen Mayhew, and Dan Roth. 2020 · 2020
Earlier work this paper cites.
Towards computational linguistics in Minangkabau language: Studies on sentiment analysis and machine translation
Fajri Koto and Ikhwan Koto. 2020 · 2020
Earlier work this paper cites.
On the language neutrality of pre-trained multilingual representations
Jindřich Libovickỳ, Rudolf Rosa, and Alexander Fraser. 2020 · 2020
Earlier work this paper cites.
CamemBERT: a tasty French language model
Louis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont, Laurent Romary, Éric de la Clergerie, Djamé Seddah, and Benoît Sagot. 2020 · 2020
Earlier work this paper cites.
JParaCrawl: A large scale web-based English-Japanese parallel corpus
Makoto Morishita, Jun Suzuki, and Masaaki Nagata. 2020 · 2020
Earlier work this paper cites.
XCOPA: A multilingual dataset for causal commonsense reasoning
Edoardo Maria Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulić, and Anna Korhonen. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Earlier work this paper cites.
AraBench: Benchmarking dialectal Arabic-English machine translation
Hassan Sajjad, Ahmed Abdelali, Nadir Durrani, and Fahim Dalvi. 2020 · 2020
Cited alongside, same era.
Bertimbau: pretrained bert models for brazilian portuguese
Fábio Souza, Rodrigo Nogueira, and Roberto Lotufo. 2020 · 2020
Cited alongside, same era.
The Tatoeba Translation Challenge – Realistic Data Sets for Low Resource and Multilingual MT
Jörg Tiedemann. 2020 · 2020
Cited alongside, same era.
SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python
Pauli Virtanen, Ralf Gommers, Travis E. Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, Stéfan J. van der Walt, Matthew Brett, Joshua Wilson, K. Jarrod Millman, Nikolay Mayorov, Andrew R. J. Nelson, Eric Jones, Robert Kern, Eric Larson, C J Carey, İlhan Polat, Yu Feng, Eric W. Moore, Jake VanderPlas, Denis Laxalde, Josef Perktold, Robert Cimrman, Ian Henriksen, E. A. Quintero, Charles R. Harris, Anne M. Archibald, Antônio H. Ribeiro, Fabian Pedregosa, Paul van Mulbregt, and SciPy 1.0 Contributors. 2020 · 2020
Cited alongside, same era.
On negative interference in multilingual models: Findings and a meta-learning treatment
Bloom library: Multimodal datasets in 300+ languages for a variety of downstream tasks
Colin Leong, Joshua Nemecek, Jacob Mansdorfer, Anna Filighera, Abraham Owodunni, and Daniel Whitenack. 2022 · 2022
Later among the works it cites.
Few-shot learning with multilingual generative language models
Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O’Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona Diab, Veselin Stoyanov, and Xian Li. 2022 · 2022
Later among the works it cites.
A balanced data approach for evaluating cross-lingual transfer: Mapping the linguistic blood bank
Dan Malkin, Tomasz Limisiewicz, and Gabriel Stanovsky. 2022 · 2022
Later among the works it cites.
TeDDi sample: Text data diversity sample for language comparison and multilingual NLP
Steven Moran, Christian Bentz, Ximena Gutierrez-Vasques, Olga Pelloni, and Tanja Samardzic. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zirui Wang, Zachary C. Lipton, and Yulia Tsvetkov. 2020 · 2020
Cited alongside, same era.
CCNet: Extracting high quality monolingual datasets from web crawl data
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave. 2020 · 2020
Cited alongside, same era.
IndoNLU: Benchmark and resources for evaluating Indonesian natural language understanding
Bryan Wilie, Karissa Vincentio, Genta Indra Winata, Samuel Cahyawijaya, Xiaohong Li, Zhi Yuan Lim, Sidik Soleman, Rahmad Mahendra, Pascale Fung, Syafri Bahar, and Ayu Purwarianti. 2020 · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Cited alongside, same era.
Are all languages created equal in multilingual BERT?
Shijie Wu and Mark Dredze. 2020 · 2020
Cited alongside, same era.
CLUE: A Chinese language understanding evaluation benchmark
Liang Xu, Hai Hu, Xuanwei Zhang, Lu Li, Chenjie Cao, Yudong Li, Yechen Xu, Kai Sun, Dian Yu, Cong Yu, Yin Tian, Qianqian Dong, Weitang Liu, Bo Shi, Yiming Cui, Junyi Li, Jun Zeng, Rongzhao Wang, Weijian Xie, Yanting Li, Yina Patterson, Zuoyu Tian, Yiwen Zhang, He Zhou, Shaoweihua Liu, Zhe Zhao, Qipeng Zhao, Cong Yue, Xinrui Zhang, Zhengliang Yang, Kyle Richardson, and Zhenzhong Lan. 2020 · 2020
Cited alongside, same era.
Alternating language modeling for cross-lingual pre-training
Jian Yang, Shuming Ma, Dongdong Zhang, Shuangzhi Wu, Zhoujun Li, and Ming Zhou. 2020 · 2020
Cited alongside, same era.
ChrEn: Cherokee-English machine translation for endangered language revitalization
Shiyue Zhang, Benjamin Frey, and Mohit Bansal. 2020 · 2020
Cited alongside, same era.
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, et al. 2022 · 2022
Later among the works it cites.
Overview of the 9th workshop on Asian translation
Toshiaki Nakazawa, Hideya Mino, Isao Goto, Raj Dabre, Shohei Higashiyama, Shantipriya Parida, Anoop Kunchukuttan, Makoto Morishita, Ondřej Bojar, Chenhui Chu, Akiko Eriguchi, Kaori Abe, Yusuke Oda, and Sadao Kurohashi. 2022 · 2022
Later among the works it cites.
An exploration of vocabulary size and transfer effects in multilingual language models for African languages
Akintunde Oladipo, Odunayo Ogundepo, Kelechi Ogueji, and Jimmy Lin. 2022 · 2022
Later among the works it cites.
Bidirectional language models are also few-shot learners
Ajay Patel, Bryan Li, Mohammad Sadegh Rasooli, Noah Constant, Colin Raffel, and Chris Callison-Burch. 2022 · 2022
Later among the works it cites.
On the role of parallel data in cross-lingual transfer learning
Machel Reid and Mikel Artetxe. 2022 · 2022
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model
Teven Le Scao, Angela Fan, Christopher Akiki, Elizabeth-Jane Pavlick, Suzana Ili’c, Daniel Hesslow, Roman Castagn’e, Alexandra Sasha Luccioni, Franccois Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Rose Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, Albert Villanova del Moral, Olatunji Ruwase, et al. 2022 · 2022
Later among the works it cites.
Towards a broad coverage named entity resource: A data-efficient approach for many diverse languages
Silvia Severini, Ayyoob Imani, Philipp Dufter, and Hinrich Schütze. 2022 · 2022
Later among the works it cites.
Aditya Siddhant, Ankur Bapna, Orhan Firat, Yuan Cao, Mia Xu Chen, Isaac Caswell, and Xavier Garcia. 2022 · 2022
Later among the works it cites.
Cree corpus: A collection of nêhiyawêwin resources
Daniela Teodorescu, Josie Matalski, Delaney Lothian, Denilson Barbosa, and Carrie Demmans Epp. 2022 · 2022
Later among the works it cites.
Cross-lingual few-shot learning on unseen languages
Genta Winata, Shijie Wu, Mayank Kulkarni, Thamar Solorio, and Daniel Preotiuc-Pietro. 2022 · 2022
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model
BigScience Workshop, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, et al. 2022 · 2022
Later among the works it cites.
Oolong: Investigating what makes crosslingual transfer hard with controlled studies
Zhengxuan Wu, Isabel Papadimitriou, and Alex Tamkin. 2022 · 2022
Later among the works it cites.
Introducing qubert: A large monolingual corpus and bert model for southern quechua
Rodolfo Zevallos, John Ortega, William Chen, Richard Castro, Nuria Bel, Cesar Toshio, Renzo Venturas, Hilario Aradiel, and Nelsi Melgarejo. 2022 · 2022
Later among the works it cites.
Ai for thai lotuscorpus
AI FOR THAI. 2023 · 2023
Later among the works it cites.
AI4Bharat
AI4Bharat. 2023 · 2023
Later among the works it cites.
Anuvaad project
Anuvaad. 2023 · 2023
Later among the works it cites.
Autshumato
Autshumato. 2023 · 2023
Later among the works it cites.
Quantifying memorization across neural language models
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. 2023 · 2023
Later among the works it cites.
Fula speech corpus
Cawoylel. 2023 · 2023
Later among the works it cites.
Cherokee corpus and Cherokee-English Dictionary
Cherokee Corpus. 2023 · 2023
Later among the works it cites.
Clarin.si
Clarin. 2023 · 2023
Later among the works it cites.
Zero-shot cross-lingual transfer language selection using linguistic similarity
Juuso Eronen, Michal Ptaszynski, and Fumito Masui. 2023 · 2023
Later among the works it cites.
Fon and french dataset
FFR Dataset. 2023 · 2023
Later among the works it cites.
Language model evaluation harness: A framework for few-shot language model evaluation
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2023 · 2023
Later among the works it cites.
Gemini: A family of highly capable multimodal models
Google DeepMind. 2023 · 2023
Later among the works it cites.
Glottolog 4.8
Harald Hammarström, Robert Forkel, Martin Haspelmath, and Sebastian Bank. 2023 · 2023
Later among the works it cites.
Machine translation benchmark dataset for languages in the horn of africa
HornMT. 2023 · 2023
Later among the works it cites.
Prompt-based methods may underestimate large language models’ linguistic generalizations
Jennifer Hu and Roger Levy. 2023 · 2023
Later among the works it cites.
Glot500: Scaling multilingual corpora and language models to 500 languages
Ayyoob Imani, Peiqin Lin, Amir Hossein Kargaran, Silvia Severini, Masoud Jalili Sabet, Nora Kassner, Chunlan Ma, Helmut Schmid, André Martins, François Yvon, and Hinrich Schütze. 2023 · 2023
Later among the works it cites.
Statistical and neural machine translation
Philipp Koehn. 2023 · 2023
Later among the works it cites.
Madlad-400: A multilingual and document-level large audited dataset
Sneha Kudugunta, Isaac Caswell, Biao Zhang, Xavier Garcia, Christopher A. Choquette-Choo, Katherine Lee, Derrick Xin, Aditya Kusupati, Romi Stella, Ankur Bapna, and Orhan Firat. 2023 · 2023
Later among the works it cites.
Lindat/clariah-cz repository
LINDAT. 2023 · 2023
Later among the works it cites.
Lingala song lyrics
Lingala Songs. 2023 · 2023
Later among the works it cites.
FinGPT: Large generative models for a small language
Risto Luukkonen, Ville Komulainen, Jouni Luoma, Anni Eskelinen, Jenna Kanerva, Hanna-Mari Kupari, Filip Ginter, Veronika Laippala, Niklas Muennighoff, Aleksandra Piktus, Thomas Wang, Nouamane Tazi, Teven Scao, Thomas Wolf, Osma Suominen, Samuli Sairanen, Mikko Merioksa, Jyrki Heinonen, Aija Vahtola, Samuel Antao, and Sampo Pyysalo. 2023 · 2023
Later among the works it cites.
Lyricstranslate
LyricsTranslate. 2023 · 2023
Later among the works it cites.
Taxi1500: A multilingual dataset for text classification in 1500 languages
Chunlan Ma, Ayyoob ImaniGooghari, Haotian Ye, Ehsaneddin Asgari, and Hinrich Schütze. 2023 · 2023
Later among the works it cites.
Umsuka isizuluparallel corpus
Rooweither Mabuya, Jade Abbott, and Vukosi Marivate. 2023 · 2023
Later among the works it cites.
Masakhane: A living collection of NLP projects for Africans, by Africans
Masakhane. 2023 · 2023
Later among the works it cites.
Scaling data-constrained language models
Niklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao, Nouamane Tazi, Aleksandra Piktus, Sampo Pyysalo, Thomas Wolf, and Colin Raffel. 2023 · 2023
Later among the works it cites.
Abkhaz text
Nart. 2023 · 2023
Later among the works it cites.
An english-kinyarwanda statistical machine translation (SMT) model
Patrick Niyongabo. 2023 · 2023
Later among the works it cites.
Mini but mighty: Efficient multilingual pretraining with linguistically-informed data selection
Tolulope Ogunremi, Dan Jurafsky, and Christopher Manning. 2023 · 2023
Later among the works it cites.
GPT-4 technical report
OpenAI. 2023 · 2023
Later among the works it cites.
Advanced Multivariate Analyses in R: Variation Partitioning
QCBS. 2023 · 2023
Later among the works it cites.
Stanford nlp group datasets
Stanford. 2023 · 2023
Later among the works it cites.
Ulukau: The Hawaiian Electronic Library
Ulukau. 2023 · 2023
Later among the works it cites.
Mingyang Wang, Heike Adel, Lukas Lange, Jannik Strötgen, and Hinrich Schütze. 2023 · 2023
Later among the works it cites.
Findings of the BabyLM challenge: Sample-efficient pretraining on developmentally plausible corpora
Alex Warstadt, Aaron Mueller, Leshem Choshen, Ethan Wilcox, Chengxu Zhuang, Juan Ciro, Rafael Mosquera, Bhargavi Paranjabe, Adina Williams, Tal Linzen, and Ryan Cotterell. 2023 · 2023
Later among the works it cites.
Wikimedia dumps
Wikimedia. 2023 · 2023
Later among the works it cites.
NusaX: Multilingual parallel sentiment dataset for 10 Indonesian local languages
Genta Indra Winata, Alham Fikri Aji, Samuel Cahyawijaya, Rahmad Mahendra, Fajri Koto, Ade Romadhony, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Pascale Fung, Timothy Baldwin, Jey Han Lau, Rico Sennrich, and Sebastian Ruder. 2023 · 2023
Later among the works it cites.
Training trajectories of language models across scales
Mengzhou Xia, Mikel Artetxe, Chunting Zhou, Xi Victoria Lin, Ramakanth Pasunuru, Danqi Chen, Luke Zettlemoyer, and Veselin Stoyanov. 2023 · 2023
Later among the works it cites.
A bit of a problem: Measurement disparities in dataset sizes across languages
Catherine Arnett, Tyler A. Chang, and Benjamin Bergen. 2024 · 2024
Closest in time.
The Belebele benchmark: A parallel reading comprehension dataset in 122 language variants
Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, and Madian Khabsa. 2024 · 2024
Closest in time.
Breaking the curse of multilinguality with cross-lingual expert language models
Terra Blevins, Tomasz Limisiewicz, Suchin Gururangan, Margaret Li, Hila Gonen, Noah A. Smith, and Luke Zettlemoyer. 2024 · 2024
Closest in time.
Goldfish may have a longer memory span than just three seconds
Lily Carey. 2024 · 2024
Closest in time.
When is multilinguality a curse? language modeling for 250 high- and low-resource languages
Tyler A. Chang, Catherine Arnett, Zhuowen Tu, and Benjamin K. Bergen. 2024a · 2024
Closest in time.
Aya Expanse: Combining research breakthroughs for a new multilingual frontier
John Dang, Shivalika Singh, Daniel D’souza, Arash Ahmadian, Alejandro Salamanca, Madeline Smith, Aidan Peppin, Sungjin Hong, Manoj Govindassamy, Terrence Zhao, Sandra Kublik, Meor Amer, Viraat Aryabumi, Jon Ander Campos, Yi-Chern Tan, Tom Kocmi, Florian Strub, Nathan Grinsztajn, Yannis Flet-Berliac, Acyr Locatelli, Hangyu Lin, Dwarak Talupuru, Bharat Venkitesh, David Cairuz, Bowen Yang, Tim Chung, Wei-Yin Ko, Sylvie Shang Shi, Amir Shukayev, Sammie Bae, Aleksandra Piktus, Roman Castagné, Felipe Cruz-Salinas, Eddie Kim, Lucas Crawhall-Stein, Adrien Morisot, Sudip Roy, Phil Blunsom, Ivan Zhang, Aidan Gomez, Nick Frosst, Marzieh Fadaee, Beyza Ermis, Ahmet Üstün, and Sara Hooker. 2024 · 2024
Closest in time.
Ethnologue, Languages of the World
Ethnologue. 2024 · 2024
Closest in time.
Mosh Levy, Alon Jacoby, and Yoav Goldberg. 2024 · 2024
Closest in time.
MaLA-500: Massive language adaptation of large language models
Peiqin Lin, Shaoxiong Ji, Jörg Tiedemann, André FT Martins, and Hinrich Schütze. 2024 · 2024
Closest in time.
Meta. 2024 · 2024
Closest in time.
Meta AI. 2024 · 2024
Closest in time.
Goldfish ® Crackers
Pepperidge Farm. 2024 · 2024
Closest in time.
Wikipedia
Wikipedia. 2024 · 2024
Closest in time.
An Analysis of Multilingual Models on Hugging Face
Catherine Arnett and Tyler Chang. 2025 · 2025
Closest in time.
Google DeepMind Gemma Team. 2025 · 2025
Closest in time.
MultiBLiMP 1.0: A massively multilingual benchmark of linguistic minimal pairs
Jaap Jumelet, Leonie Weissweiler, and Arianna Bisazza. 2025 · 2025
Closest in time.
Fineweb2: One pipeline to scale them all — adapting pre-training data processing to every language
Guilherme Penedo, Hynek Kydlíček, Vinko Sabolčec, Bettina Messmer, Negar Foroutan, Amir Hossein Kargaran, Colin Raffel, Martin Jaggi, Leandro Von Werra, and Thomas Wolf. 2025 · 2025
Closest in time.
Team Gemma, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, Morgane Rivière, et al. 2025 · 2025
Closest in time.
Commonlid: Re-evaluating state-of-the-art language identification performance on web data
Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett, Rafael Mosquera-Gómez, Sara Hincapie-Monsalve, Thom Vaughan, Damian Stewart, Malte Ostendorff, Idris Abdulmumin, Vukosi Marivate, et al. 2026 · 2026
Closest in time.
Multilingual open text release 1: Public domain news in 44 languages
Chester Palen-Michel, June Kim, and Constantine Lignos. 2022 · 2089
Closest in time.