Fetching the paper…
Reading the bibliography…
Due to the lack of quality data for low-resource Bantu languages, significant challenges are presented in text classification and other practical implementations.
Eda: Easy data augmentation techniques for boosting performance on text classification tasks
Jason Wei and Kai Zou. 2019 · 1901
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2023 · 1910
Earlier work this paper cites.
WordNet: A lexical database for English
George A. Miller. 1994 · 1994
Earlier work this paper cites.
Jiaao Chen, Zichao Yang, and Diyi Yang. 2020 · 2004
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Improving neural machine translation models with monolingual data
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang. 2017 · 2017
Earlier work this paper cites.
Seqgan: Sequence generative adversarial nets with policy gradient
Lantao Yu, Weinan Zhang, Jun Wang, and Yong Yu. 2017 · 2017
Earlier work this paper cites.
Contextual augmentation: Data augmentation by words with paradigmatic relations
Sosuke Kobayashi. 2018 · 2018
Earlier work this paper cites.
Unsupervised machine translation using monolingual corpora only
Guillaume Lample, Alexis Conneau, Ludovic Denoyer, and Marc’Aurelio Ranzato. 2018 · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
AEDA: An easier data augmentation technique for text classification
Akbar Karimi, Leonardo Rossi, and Andrea Prati. 2021 · 2021
Cited alongside, same era.
Small data? no problem! exploring the viability of pretrained multilingual language models for low-resourced languages
Kelechi Ogueji, Yuxin Zhu, and Jimmy Lin. 2021 · 2021
Cited alongside, same era.
Natural Language Processing for African Languages
David Ifeoluwa Adelani. 2022 · 2022
Cited alongside, same era.
Adapting pre-trained language models to African languages via multilingual adaptive fine-tuning
Jesujoba O. Alabi, David Ifeoluwa Adelani, Marius Mosbach, and Dietrich Klakow. 2022 · 2022
Cited alongside, same era.
Zhuoyan Li, Hangxiao Zhu, Zhuoran Lu, and Ming Yin. 2023 · 2023
Later among the works it cites.
SemEval-2023 Task 12: Sentiment Analysis for African Languages (AfriSenti-SemEval)
Shamsuddeen Hassan Muhammad, Idris Abdulmumin, Seid Muhie Yimam, David Ifeoluwa Adelani, Ibrahim Sa’id Ahmad, Nedjma Ousidhoum, Abinew Ali Ayele, Saif M. Mohammad, Meriem Beloucif, and Sebastian Ruder. 2023 · 2023
Later among the works it cites.
Cross-lingual retrieval augmented prompt for low-resource languages
Ercong Nie, Sheng Liang, Helmut Schmid, and Hinrich Schütze. 2023 · 2023
Later among the works it cites.
Bantuberta: Using language family grouping in multilingual language modeling for bantu languages
Jesse Parvess. 2023 · 2023
Later among the works it cites.
Text augmentation using dataset reconstruction for low-resource classification
Adir Rahamim, Guy Uziel, Esther Goldbraich, and Ateret Anaby Tavor. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Md Saroar Jahan, Djamila Romaissa Beddiar, Mourad Oussalah, and Muhidin Mohamed. 2022 · 2022
Cited alongside, same era.
Combining WordNet and word embeddings in data augmentation for legal texts
Sezen Perçin, Andrea Galassi, Francesca Lagioia, Federico Ruggeri, Piera Santin, Giovanni Sartor, and Paolo Torroni. 2022 · 2022
Cited alongside, same era.
Data augmentation for intent classification with off-the-shelf large language models
Gaurav Sahu, Pau Rodriguez, Issam Laradji, Parmida Atighehchian, David Vazquez, and Dzmitry Bahdanau. 2022 · 2022
Cited alongside, same era.
EPiDA: An easy plug-in data augmentation framework for high performance text classification
Minyi Zhao, Lu Zhang, Yi Xu, Jiandong Ding, Jihong Guan, and Shuigeng Zhou. 2022 · 2022
Cited alongside, same era.
To augment or not to augment? a comparative study on text augmentation techniques for low-resource nlp
Gözde Gül Şahin. 2022 · 2022
Cited alongside, same era.
Later among the works it cites.
Lida: Language-independent data augmentation for text classification
Yudianto Sujana and Hung-Yu Kao. 2023 · 2023
Later among the works it cites.
State of nlp in kenya: A survey
Cynthia Jayne Amol, Everlyn Asiko Chimoto, Rose Delilah Gesicho, Antony M. Gitau, Naome A. Etori, Caringtone Kinyanjui, Steven Ndung’u, Lawrence Moruye, Samson Otieno Ooko, Kavengi Kitonga, Brian Muhia, Catherine Gitau, Antony Ndolo, Lilian D. A. Wanzare, Albert Njoroge Kahira, and Ronald Tombe. 2024 · 2024
Later among the works it cites.
Inditext boost: Text augmentation for low resource india languages
Onkar Litake, Niraj Yagnik, and Shreyas Rajesh Labhsetwar. 2024 · 2024
Later among the works it cites.
Cross-lingual transfer of multilingual models on low resource african languages
Harish Thangaraj, Ananya Chenat, Jaskaran Singh Walia, and Vukosi Marivate. 2024 · 2024
Later among the works it cites.