Fetching the paper…
Reading the bibliography…
The disparity in the languages commonly studied in Natural Language Processing (NLP) is typically reflected by referring to languages as low vs high-resourced.
Bridging the gap between language models and cross-lingual sequence labeling
Nuo Chen, Linjun Shou, Ming Gong, Jian Pei, and Daxin Jiang. 2022 · 1923
Earlier work this paper cites.
Enhancing cross-lingual natural language inference by prompt-learning from cross-lingual templates
Kunxun Qi, Hai Wan, Jianfeng Du, and Haolan Chen. 2022 · 1923
Earlier work this paper cites.
The makerere radio speech corpus: A Luganda radio corpus for automatic speech recognition
Jonathan Mukiibi, Andrew Katumba, Joyce Nakatumba-Nabende, Ali Hussein, and Joshua Meyer. 2022 · 1954
Earlier work this paper cites.
A history of the hunting peoples of the northern east africa coast: Ecological and socio-economic considerations
Daniel Stiles. 1982 · 1982
Earlier work this paper cites.
UXLA: A robust unsupervised data augmentation framework for zero-resource cross-lingual NLP
M Saiful Bari, Tasnim Mohiuddin, and Shafiq Joty. 2021 · 1992
Earlier work this paper cites.
Dahalo: an Endangered Language
Mauro Tosco. 1992 · 1992
Earlier work this paper cites.
Kirrkirr: Software for Browsing and Visual Exploration of a Structured Warlpiri Dictionary
Christopher D. Manning, Kevin Jansz, and Nitin Indurkhya. 2001 · 2001
Earlier work this paper cites.
UNESCO Ad Hoc Expert Group on Endangered Languages
Matthias Brenzinger, Arienne M. Dwyer, Tjeerd de Graaf, Colette Grinevald, Michael Krauss, Osahito Miyaoka, Nicholas Ostler, Osamu Sakiyama, María E. Villalón, Akira Y. Yamamoto, and Ofelia Zepeda. 2003 · 2003
Earlier work this paper cites.
Using thematic analysis in psychology
Virginia Braun and Victoria Clarke. 2006 · 2006
Earlier work this paper cites.
Human Language Technology Resources for Less Commonly Taught Languages: Lessons Learned Toward Creation of Basic Language Resources
Heather Simpson, Christopher Cieri, Kazuaki Maeda, Kathryn Baker, and Boyan Onyshkevych. 2008 · 2008
Earlier work this paper cites.
. a survey of computational morphological resources for low-density languages
H Hammarström. 2009 · 2009
Earlier work this paper cites.
Portraits of language vitality in the languages of Indonesia , pages 19–47
Karl Anderbeck. 2015 · 2015
Earlier work this paper cites.
The ethical limits of bungee research in ICTD
Andy Dearden and William D. Tucker. 2021 · 2015
Earlier work this paper cites.
Selection Criteria for Low Resource Language Programs
Christopher Cieri, Mike Maxwell, Stephanie Strassel, and Jennifer Tracey. 2016 · 2016
Earlier work this paper cites.
Aboriginal world views and colonisation: implications for coastal sustainability†
Leonard Collard Laura Stocker and Angela Rooney. 2016 · 2016
Earlier work this paper cites.
Model transfer for tagging low-resource languages using a bilingual dictionary
Meng Fang and Trevor Cohn. 2017 · 2017
Earlier work this paper cites.
Cross-lingual dependency parsing with late decoding for truly low-resource languages
Michael Schlichtkrull and Anders Søgaard. 2017 · 2017
Earlier work this paper cites.
Towards automating healthcare question answering in a noisy multilingual low-resource setting
Jeanne E. Daniel, Willie Brink, Ryan Eloff, and Charles Copley. 2019 · 2019
Earlier work this paper cites.
Generalized data augmentation for low-resource translation
Mengzhou Xia, Xiang Kong, Antonios Anastasopoulos, and Graham Neubig. 2019 · 2019
Earlier work this paper cites.
Improving low-resource cross-lingual document retrieval by reranking with deep bilingual representations
Rui Zhang, Caitlin Westerfield, Sungrok Shim, Garrett Bingham, Alexander Fabbri, William Hu, Neha Verma, and Dragomir Radev. 2019 · 2019
Earlier work this paper cites.
Identifying sentiments in Algerian code-switched user-generated comments
Wafia Adouane, Samia Touileb, and Jean-Philippe Bernardy. 2020 · 2020
Earlier work this paper cites.
Massive vs. curated embeddings for low-resourced languages: the case of Yorùbá and Twi
Jesujoba Alabi, Kwabena Amponsah-Kaakyire, David Adelani, and Cristina España-Bonet. 2020 · 2020
Earlier work this paper cites.
Towards computational resource grammars for Runyankore and rukiga
David Bamutura, Peter Ljunglöf, and Peter Nebende. 2020 · 2020
Earlier work this paper cites.
Semi-supervised development of ASR systems for multilingual code-switched speech in under-resourced languages
Astik Biswas, Emre Yilmaz, Febe De Wet, Ewald Van der westhuizen, and Thomas Niesler. 2020 · 2020
Earlier work this paper cites.
Entity Linking in 100 Languages
Jan A. Botha, Zifei Shan, and Daniel Gillick. 2020 · 2020
Earlier work this paper cites.
Exploring a Choctaw language corpus with word vectors and minimum distance length
Jacqueline Brixey, David Sides, Timothy Vizthum, David Traum, and Khalil Iskarous. 2020 · 2020
Earlier work this paper cites.
No data to crawl? monolingual corpus creation from PDF files of truly low-resource languages in Peru
Gina Bustamante, Arturo Oncevay, and Roberto Zariquiey. 2020 · 2020
Earlier work this paper cites.
Data Augmentation via Subtree Swapping for Dependency Parsing of Low-Resource Languages
Mathieu Dehouck and Carlos Gómez-Rodríguez. 2020 · 2020
Earlier work this paper cites.
DAGA: Data augmentation with a generation approach for low-resource tagging tasks
Bosheng Ding, Linlin Liu, Lidong Bing, Canasai Kruengkrai, Thien Hai Nguyen, Shafiq Joty, Luo Si, and Chunyan Miao. 2020 · 2020
Earlier work this paper cites.
Unsupervised cross-lingual part-of-speech tagging for truly low-resource scenarios
Ramy Eskander, Smaranda Muresan, and Michael Collins. 2020b · 2020
Earlier work this paper cites.
Cross-lingual unsupervised sentiment classification with multi-view transfer learning
Hongliang Fei and Ping Li. 2020 · 2020
Earlier work this paper cites.
Neural machine translation models with back-translation for the extremely low-resource indigenous language Bribri
Isaac Feldman and Rolando Coto-Solano. 2020 · 2020
Earlier work this paper cites.
Processing language resources of under-resourced and endangered languages for the generation of augmentative alternative communication boards
Anne Ferger. 2020 · 2020
Earlier work this paper cites.
Not low-resource anymore: Aligner ensembling, batch filtering, and new datasets for Bengali-English machine translation
Tahmid Hasan, Abhik Bhattacharjee, Kazi Samin, Masum Hasan, Madhusudan Basak, M. Sohel Rahman, and Rifat Shahriyar. 2020 · 2020
Earlier work this paper cites.
Unsupervised morphological paradigm completion
Huiming Jin, Liwei Cai, Yihui Peng, Chen Xia, Arya McCarthy, and Katharina Kann. 2020 · 2020
Earlier work this paper cites.
The State and Fate of Linguistic Diversity and Inclusion in the NLP World
Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020 · 2020
Earlier work this paper cites.
Simulated multiple reference training improves low-resource machine translation
Huda Khayrallah, Brian Thompson, Matt Post, and Philipp Koehn. 2020 · 2020
Earlier work this paper cites.
Interactive word completion for morphologically complex languages
William Lane and Steven Bird. 2020 · 2020
Earlier work this paper cites.
Towards instance-level parser selection for cross-lingual transfer of dependency parsers
Robert Litschko, Ivan Vulić, Željko Agić, and Goran Glavaš. 2020 · 2020
Earlier work this paper cites.
Analogy models for neural word inflection
Ling Liu and Mans Hulden. 2020 · 2020
Earlier work this paper cites.
Tackling the low-resource challenge for canonical segmentation
Manuel Mager, Özlem Çetinoğlu, and Katharina Kann. 2020 · 2020
Earlier work this paper cites.
Learnings from technological interventions in a low resource language: A case-study on Gondi
Devansh Mehta, Sebastin Santy, Ramaravind Kommiya Mothilal, Brij Mohan Lal Srivastava, Alok Sharma, Anurag Shukla, Vishnu Prasad, Venkanna U, Amit Sharma, and Kalika Bali. 2020 · 2020
Earlier work this paper cites.
An analysis of massively multilingual neural machine translation for low-resource languages
Aaron Mueller, Garrett Nicolai, Arya D. McCarthy, Dylan Lewis, Winston Wu, and David Yarowsky. 2020 · 2020
Earlier work this paper cites.
Participatory Research for Low-resourced Machine Translation: A Case Study in African Languages
Wilhelmina Nekoto, Vukosi Marivate, Tshinondiwa Matsila, Timi Fasubaa, Taiwo Fagbohungbe, Solomon Oluwole Akinola, Shamsuddeen Muhammad, Salomon Kabongo Kabenamualu, Salomey Osei, Freshia Sackey, Rubungo Andre Niyongabo, Ricky Macharm, Perez Ogayo, Orevaoghene Ahia, Musie Meressa Berhe, Mofetoluwa Adeyemi, Masabata Mokgesi-Selinga, Lawrence Okegbemi, Laura Martinus, Kolawole Tajudeen, Kevin Degila, Kelechi Ogueji, Kathleen Siminyu, Julia Kreutzer, Jason Webster, Jamiil Toure Ali, Jade Abbott, Iroro Orife, Ignatius Ezeani, Idris Abdulkadir Dangana, Herman Kamper, Hady Elsahar, Goodness Duru, Ghollah Kioko, Murhabazi Espoir, Elan van Biljon, Daniel Whitenack, Christopher Onyefuluchi, Chris Chinenye Emezue, Bonaventure F. P. Dossou, Blessing Sibanda, Blessing Bassey, Ayodele Olabiyi, Arshath Ramkilowan, Alp Öktem, Adewale Akinfaderin, and Abdallah Bashir. 2020 · 2020
Earlier work this paper cites.
KINNEWS and KIRNEWS: Benchmarking cross-lingual text classification for Kinyarwanda and Kirundi
Rubungo Andre Niyongabo, Qu Hong, Julia Kreutzer, and Li Huang. 2020 · 2020
Earlier work this paper cites.
MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual Transfer
Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych, and Sebastian Ruder. 2020 · 2020
Earlier work this paper cites.
SimplifyUR: Unsupervised lexical text simplification for Urdu
Namoos Hayat Qasmi, Haris Bin Zia, Awais Athar, and Agha Ali Raza. 2020 · 2020
Earlier work this paper cites.
A survey of konkani nlp resources
Annie Rajan, Ambuja Salgaonkar, and Ramprasad Joshi. 2020 · 2020
Earlier work this paper cites.
Cross-lingual emotion lexicon induction using representation alignment in low-resource settings
Arun Ramachandran and Gerard de Melo. 2020 · 2020
Earlier work this paper cites.
Soft gazetteers for low-resource named entity recognition
Shruti Rijhwani, Shuyan Zhou, Graham Neubig, and Jaime Carbonell. 2020 · 2020
Earlier work this paper cites.
Identification of indigenous knowledge concepts through semantic networks, spelling tools and word embeddings
Renato Rocha Souza, Amelie Dorn, Barbara Piringer, and Eveline Wandl-Vogt. 2020 · 2020
Earlier work this paper cites.
Detecting urgency status of crisis tweets: A transfer learning approach for low resource languages
Efsun Sarioglu Kayi, Linyong Nan, Bohan Qu, Mona Diab, and Kathleen McKeown. 2020 · 2020
Earlier work this paper cites.
Analysing cross-lingual transfer in lemmatisation for Indian languages
Kumar Saurav, Kumar Saunack, and Pushpak Bhattacharyya. 2020 · 2020
Earlier work this paper cites.
Leveraging monolingual data with self-supervision for multilingual neural machine translation
Aditya Siddhant, Ankur Bapna, Yuan Cao, Orhan Firat, Mia Chen, Sneha Kudugunta, Naveen Arivazhagan, and Yonghui Wu. 2020 · 2020
Cited alongside, same era.
CPLM, a parallel corpus for Mexican languages: Development and interface
Gerardo Sierra Martínez, Cynthia Montaño, Gemma Bel-Enguix, Diego Córdova, and Margarita Mota Montoya. 2020 · 2020
Cited alongside, same era.
Getting more data for low-resource morphological inflection: Language models and data augmentation
Alexey Sorokin. 2020 · 2020
Cited alongside, same era.
UDapter: Language adaptation for truly Universal Dependency parsing
Ahmet Üstün, Arianna Bisazza, Gosse Bouma, and Gertjan van Noord. 2020 · 2020
Cited alongside, same era.
Structure-level knowledge distillation for multilingual sequence labeling
Xinyu Wang, Yong Jiang, Nguyen Bach, Tao Wang, Fei Huang, and Kewei Tu. 2020 · 2020
Cited alongside, same era.
Out of thin air: Is zero-shot cross-lingual keyword detection better than unsupervised?
Boshko Koloski, Senja Pollak, Blaž Škrlj, and Matej Martinc. 2022 · 2022
Later among the works it cites.
Meta-learning for fast cross-lingual adaptation in dependency parsing
Anna Langedijk, Verna Dankers, Phillip Lippe, Sander Bos, Bryan Cardenas Guevara, Helen Yannakoudakis, and Ekaterina Shutova. 2022 · 2022
Later among the works it cites.
Bloom library: Multimodal datasets in 300+ languages for a variety of downstream tasks
Colin Leong, Joshua Nemecek, Jacob Mansdorfer, Anna Filighera, Abraham Owodunni, and Daniel Whitenack. 2022 · 2022
Later among the works it cites.
Multi-level distillation of semantic knowledge for pre-training multilingual language model
Mingqi Li, Fei Ding, Dan Zhang, Long Cheng, Hongxin Hu, and Feng Luo. 2022a · 2022
Later among the works it cites.
Low resource style transfer via domain adaptive meta learning
Xiangyang Li, Xiang Long, Yu Xia, and Sujian Li. 2022b · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Exploring Amharic sentiment analysis from social media texts: Building annotation tools and classification models
Seid Muhie Yimam, Hizkiel Mitiku Alemayehu, Abinew Ayele, and Chris Biemann. 2020 · 2020
Cited alongside, same era.
Interactive refinement of cross-lingual word embeddings
Michelle Yuan, Mozhi Zhang, Benjamin Van Durme, Leah Findlater, and Jordan Boyd-Graber. 2020 · 2020
Cited alongside, same era.
Improving candidate generation for low-resource cross-lingual entity linking
Shuyan Zhou, Shruti Rijhwani, John Wieting, Jaime Carbonell, and Graham Neubig. 2020 · 2020
Cited alongside, same era.
Long document summarization in a low resource setting using pretrained language models
Ahsaas Bajaj, Pavitra Dangati, Kalpesh Krishna, Pradhiksha Ashok Kumar, Rheeya Uppaal, Goldman Sachs, Bradford Windsor, Eliot Brenner, Dominic Dotterrer, Rajarshi Das, et al. 2021 · 2021
Cited alongside, same era.
Reducing confusion in active learning for part-of-speech tagging
Aditi Chaudhary, Antonios Anastasopoulos, Zaid Sheikh, and Graham Neubig. 2021 · 2021
Cited alongside, same era.
Towards more equitable question answering systems: How much more data do you need?
Arnab Debnath, Navid Rajabi, Fardina Fathmiul Alam, and Antonios Anastasopoulos. 2021 · 2021
Cited alongside, same era.
Few shot dialogue state tracking using meta-learning
Saket Dingliwal, Shuyang Gao, Sanchit Agarwal, Chien-Wei Lin, Tagyoung Chung, and Dilek Hakkani-Tur. 2021 · 2021
Cited alongside, same era.
Label-aware multi-level contrastive learning for cross-lingual spoken language understanding
Shining Liang, Linjun Shou, Jian Pei, Ming Gong, Wanli Zuo, Xianglin Zuo, and Daxin Jiang. 2022 · 2022
Later among the works it cites.
Conversational speech recognition needs data? experiments with Austrian German
Julian Linke, Philip N. Garner, Gernot Kubin, and Barbara Schuppler. 2022 · 2022
Later among the works it cites.
Language-agnostic meta-learning for low-resource text-to-speech with articulatory features
Florian Lux and Thang Vu. 2022 · 2022
Later among the works it cites.
Bilingual lexicon induction for low-resource languages using graph matching via optimal transport
Kelly Marchisio, Ali Saad-Eldin, Kevin Duh, Carey Priebe, and Philipp Koehn. 2022 · 2022
Later among the works it cites.
WordNet-QU: Development of a lexical database for Quechua varieties
Nelsi Melgarejo, Rodolfo Zevallos, Hector Gomez, and John E. Ortega. 2022 · 2022
Later among the works it cites.
ViHealthBERT: Pre-trained language models for Vietnamese in health text mining
Nguyen Minh, Vu Hoang Tran, Vu Hoang, Huy Duc Ta, Trung Huu Bui, and Steven Quoc Hung Truong. 2022 · 2022
Later among the works it cites.
WECHSEL: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language models
Benjamin Minixhofer, Fabian Paischer, and Navid Rekabsaz. 2022 · 2022
Later among the works it cites.
SHONGLAP: A large Bengali open-domain dialogue corpus
Syed Mostofa Monsur, Sakib Chowdhury, Md Shahrar Fatemi, and Shafayat Ahmed. 2022 · 2022
Later among the works it cites.
Eeny, meeny, miny, moe. how to choose data for morphological inflection
Saliha Muradoglu and Mans Hulden. 2022 · 2022
Later among the works it cites.
KinyaBERT: a morphology-aware Kinyarwanda language model
Antoine Nzeyimana and Andre Niyongabo Rubungo. 2022 · 2022
Later among the works it cites.
An inflectional database for gitksan
Bruce Oliver, Clarissa Forbes, Changbing Yang, Farhan Samir, Edith Coates, Garrett Nicolai, and Miikka Silfverberg. 2022 · 2022
Later among the works it cites.
BAD-X: Bilingual adapters improve zero-shot cross-lingual transfer
Marinela Parović, Goran Glavaš, Ivan Vulić, and Anna Korhonen. 2022 · 2022
Later among the works it cites.
AsNER - annotated dataset and baseline for Assamese named entity recognition
Dhrubajyoti Pathak, Sukumar Nandi, and Priyankoo Sarmah. 2022 · 2022
Later among the works it cites.
Overlap-based vocabulary generation improves cross-lingual transfer among related languages
Vaidehi Patil, Partha Talukdar, and Sunita Sarawagi. 2022 · 2022
Later among the works it cites.
Data-efficient strategies for expanding hate speech detection into under-resourced languages
Paul Röttger, Debora Nozza, Federico Bianchi, and Dirk Hovy. 2022 · 2022
Later among the works it cites.
Don’t stop fine-tuning: On training regimes for few-shot cross-lingual transfer with multilingual language models
Fabian David Schmidt, Ivan Vulić, and Goran Glavaš. 2022 · 2022
Later among the works it cites.
Primum Non Nocere: Before working with Indigenous data, the ACL must confront ongoing colonialism
Lane Schwartz. 2022 · 2022
Later among the works it cites.
HAWP: a dataset for Hindi arithmetic word problem solving
Harshita Sharma, Pruthwik Mishra, and Dipti Sharma. 2022 · 2022
Later among the works it cites.
BembaSpeech: A speech recognition corpus for the Bemba language
Claytone Sikasote and Antonios Anastasopoulos. 2022 · 2022
Later among the works it cites.
Assessing digital language support on a global scale
Gary F. Simons, Abbey L. L. Thomas, and Chad K. K. White. 2022 · 2022
Later among the works it cites.
Multi-task pre-training for plug-and-play task-oriented dialogue system
Yixuan Su, Lei Shu, Elman Mansimov, Arshit Gupta, Deng Cai, Yi-An Lai, and Yi Zhang. 2022 · 2022
Later among the works it cites.
Language branch gated multilingual neural machine translation
Haoran Sun and Deyi Xiong. 2022 · 2022
Later among the works it cites.
Incorporating LIWC in neural networks to improve human trait and behavior analysis in low resource scenarios
Isil Yakut Kilic and Shimei Pan. 2022 · 2022
Later among the works it cites.
Language ideologies, policies and practices within the multilingual Kenyan context
David Barasa. 2023 · 2023
Later among the works it cites.
Making more of little data: Improving low-resource automatic speech recognition using data augmentation
Martijn Bartelds, Nay San, Bradley McDonnell, Dan Jurafsky, and Martijn Wieling. 2023 · 2023
Later among the works it cites.
Adversarial training for low-resource disfluency correction
Vineet Bhat, Preethi Jyothi, and Pushpak Bhattacharyya. 2023 · 2023
Later among the works it cites.
Generative AI has a language problem
Monojit Choudhury. 2023 · 2023
Later among the works it cites.
Meeting the needs of low-resource languages: The value of automatic alignments via pretrained models
Abteen Ebrahimi, Arya D. McCarthy, Arturo Oncevay, John E. Ortega, Luis Chiruzzo, Gustavo Giménez-Lugo, Rolando Coto-Solano, and Katharina Kann. 2023 · 2023
Later among the works it cites.
Question-answering in a low-resourced language: Benchmark dataset and models for Tigrinya
Fitsum Gaim, Wonsuk Yang, Hancheol Park, and Jong Park. 2023 · 2023
Later among the works it cites.
DALE: Generative data augmentation for low-resource legal NLP
Sreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar, S Ramaneswaran, S Sakshi, Utkarsh Tyagi, and Dinesh Manocha. 2023 · 2023
Later among the works it cites.
A Material Lens on Coloniality in NLP
William Held, Camille Harris, Michael Best, and Diyi Yang. 2023 · 2023
Later among the works it cites.
Improving long dialogue summarization with semantic graph representation
Yilun Hua, Zhaoyuan Deng, and Kathleen McKeown. 2023 · 2023
Later among the works it cites.
Glot500: Scaling multilingual corpora and language models to 500 languages
Ayyoob ImaniGooghari, Peiqin Lin, Amir Hossein Kargaran, Silvia Severini, Masoud Jalili Sabet, Nora Kassner, Chunlan Ma, Helmut Schmid, André Martins, François Yvon, and Hinrich Schütze. 2023 · 2023
Later among the works it cites.
Node placement in argument maps: Modeling unidirectional relations in high & low-resource scenarios
Iman Jundi, Neele Falk, Eva Maria Vecchi, and Gabriella Lapesa. 2023 · 2023
Later among the works it cites.
The semantic scholar open data platform
Rodney Kinney, Chloe Anastasiades, Russell Authur, Iz Beltagy, Jonathan Bragg, Alexandra Buraczynski, Isabel Cachola, Stefan Candra, Yoganand Chandrasekhar, Arman Cohan, et al. 2023 · 2023
Later among the works it cites.
Multijugate dual learning for low-resource task-oriented dialogue system
Shimin Li, Xiaotian Zhang, Yanjun Zheng, Linyang Li, and Xipeng Qiu. 2023a · 2023
Later among the works it cites.
ViT-TTS: Visual text-to-speech with scalable diffusion transformer
Huadai Liu, Rongjie Huang, Xuan Lin, Wenqiang Xu, Maozong Zheng, Hong Chen, Jinzheng He, and Zhou Zhao. 2023a · 2023
Later among the works it cites.
Small data, big impact: Leveraging minimal data for effective machine translation
Jean Maillard, Cynthia Gao, Elahe Kalbassi, Kaushik Ram Sadagopan, Vedanuj Goswami, Philipp Koehn, Angela Fan, and Francisco Guzman. 2023 · 2023
Later among the works it cites.
Multi3NLU++: A multilingual, multi-intent, multi-domain dataset for natural language understanding in task-oriented dialogue
Nikita Moghe, Evgeniia Razumovskaia, Liane Guillou, Ivan Vulić, Anna Korhonen, and Alexandra Birch. 2023 · 2023
Later among the works it cites.
Unsupervised graph-text mutual conversion with a unified pretrained language model
Yi Xu, Shuqian Sheng, Jiexing Qi, Luoyi Fu, Zhouhan Lin, Xinbing Wang, and Chenghu Zhou. 2023 · 2023
Later among the works it cites.
Soft language clustering for multilingual model pre-training
Jiali Zeng, Yufan Jiang, Yongjing Yin, Yi Jing, Fandong Meng, Binghuai Lin, Yunbo Cao, and Jie Zhou. 2023 · 2023
Later among the works it cites.
The Indigenous Languages Technology (ILT) project at the National Research Council of Canada, and its context - NRC Publications Archive - Canada.ca
Roland Kuhn. 2024 · 2024
Closest in time.
“I Searched for a Religious Song in Amharic and Got Sexual Content Instead”: Investigating Online Harm in Low-Resourced Languages on YouTube
Hellina Hailu Nigatu and Inioluwa Deborah Raji. 2024 · 2024
Closest in time.
Scaling neural machine translation to 200 languages
NLLB. 2024 · 2024
Closest in time.
Aya model: An instruction finetuned open-access multilingual language model
Ahmet Üstün, Viraat Aryabumi, Zheng-Xin Yong, Wei-Yin Ko, Daniel D’souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, et al. 2024 · 2024
Closest in time.
Local word discovery for interactive transcription
William Lane and Steven Bird. 2021a · 2067
Closest in time.
Local Word Discovery for Interactive Transcription
William Lane and Steven Bird. 2021b · 2067
Closest in time.