Fetching the paper…
Reading the bibliography…
How can large language models (LLMs) process and translate endangered languages? Many languages lack a large corpus to train a decent LLM; therefore existing LLMs rarely perform well in unseen, endangered languages.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Notes on wolof grammar by william a. stewart adapted for the present text by william w. gage
William A. Stewart. 1970 · 1970
Earlier work this paper cites.
Gitksan Grammar
Bruce Rigsby. 1986 · 1986
Earlier work this paper cites.
Wollof - english dictionary
Peace Corps The Gambia. 1995 · 1995
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
Manchu grammar
Liliya M Gorelova. 2002 · 2002
Earlier work this paper cites.
The Arapaho Language
Andrew Cowell and Alonzo Moss. 2008 · 2008
Earlier work this paper cites.
Foma: a finite-state compiler and library
Mans Hulden. 2009 · 2009
Earlier work this paper cites.
Atlas of the World’s Languages in Danger
Christopher Moseley. 2010 · 2010
Earlier work this paper cites.
Glottolog/langdoc: Defining dialects, languages, and language families as collections of resources
Sebastian Nordhoff and Harald Hammarström. 2011 · 2011
Earlier work this paper cites.
Creating lexical resources for polysynthetic languages—the case of Arapaho
Ghazaleh Kazeminejad, Andrew Cowell, and Mans Hulden. 2017 · 2017
Earlier work this paper cites.
Gramática de la lengua bribri
C.V. Jara. 2018 · 2018
Earlier work this paper cites.
A neural morphological analyzer for Arapaho verbs learned from a finite state transducer
Sarah Moeller, Ghazaleh Kazeminejad, Andrew Cowell, and Mans Hulden. 2018 · 2018
Earlier work this paper cites.
The modeling of bribri verbal morphology
Sofía Flores-Solórzano. 2019 · 2019
Cited alongside, same era.
A comprehensive Manchu-English dictionary
Jerry Norman. 2020 · 2020
Cited alongside, same era.
Grammaire du Nalögo, langue océanienne de l’île Santa Cruz (Archipel des îles Salomon)
Valentina Alfarano. 2021 · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021 · 2021
Cited alongside, same era.
An FST morphological analyzer for the gitksan language
Clarissa Forbes, Garrett Nicolai, and Miikka Silfverberg. 2021 · 2021
Cited alongside, same era.
Building machine translation systems for the next thousand languages
How to design translation prompts for chatgpt: An empirical study
Yuan Gao, Ruili Wang, and Feng Hou. 2023 · 2023
Later among the works it cites.
Findings of the SIGMORPHON 2023 shared task on interlinear glossing
Michael Ginn, Sarah Moeller, Alexis Palmer, Anna Stacey, Garrett Nicolai, Mans Hulden, and Miikka Silfverberg. 2023 · 2023
Later among the works it cites.
How good are gpt models at machine translation? a comprehensive evaluation
Amr Hendy, Mohamed Abdelrehim, Amr Sharaf, Vikas Raunak, Mohamed Gabr, Hitokazu Matsushita, Young Jin Kim, Mohamed Afify, and Hany Hassan Awadalla. 2023 · 2023
Later among the works it cites.
Haoyang Huang, Tianyi Tang, Dongdong Zhang, Wayne Xin Zhao, Ting Song, Yan Xia, and Furu Wei. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ankur Bapna, Isaac Caswell, Julia Kreutzer, Orhan Firat, Daan van Esch, Aditya Siddhant, Mengmeng Niu, Pallavi Baljekar, Xavier Garcia, Wolfgang Macherey, Theresa Breiner, Vera Axelrod, Jason Riesa, Yuan Cao, Mia Xu Chen, Klaus Macherey, Maxim Krikun, Pidong Wang, Alexander Gutkin, Apurva Shah, Yanping Huang, Zhifeng Chen, Yonghui Wu, and Macduff Hughes. 2022 · 2022
Cited alongside, same era.
The Flores-101 evaluation benchmark for low-resource and multilingual machine translation
Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzmán, and Angela Fan. 2022 · 2022
Cited alongside, same era.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Cited alongside, same era.
Language models are multilingual chain-of-thought reasoners
Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, et al. 2022 · 2022
Cited alongside, same era.
No language left behind: Scaling human-centered machine translation
NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang. 2022 · 2022
Cited alongside, same era.
Mega: Multilingual evaluation of generative ai
Kabir Ahuja, Rishav Hada, Millicent Ochieng, Prachi Jain, Harshita Diddee, Samuel Maina, Tanuja Ganu, Sameer Segal, Maxamed Axmed, Kalika Bali, et al. 2023 · 2023
Cited alongside, same era.
Wenxiang Jiao, Wenxuan Wang, JT Huang, Xing Wang, and ZP Tu. 2023 · 2023
Later among the works it cites.
Bribri–spanish spanish–bribri dictionary
Krohn, H. S. 2023 · 2023
Later among the works it cites.
Chatgpt mt: Competitive for high-(but not low-) resource languages
Nathaniel R Robinson, Perez Ogayo, David R Mortensen, and Graham Neubig. 2023 · 2023
Later among the works it cites.
A benchmark for learning to translate a new language from one grammar book
Garrett Tanzer, Mirac Suzgun, Eline Visser, Dan Jurafsky, and Luke Melas-Kyriazi. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Later among the works it cites.
Lego-MT: Learning detachable models for massively multilingual machine translation
Fei Yuan, Yinquan Lu, Wenhao Zhu, Lingpeng Kong, Lei Li, Yu Qiao, and Jingjing Xu. 2023 · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023 · 2023
Later among the works it cites.
Multilingual machine translation with large language models: Empirical results and analysis
Wenhao Zhu, Hongyi Liu, Qingxiu Dong, Jingjing Xu, Lingpeng Kong, Jiajun Chen, Lei Li, and Shujian Huang. 2024 · 2024
Closest in time.