Fetching the paper…
Reading the bibliography…
Linguistic disparity in the NLP world is a problem that has been widely acknowledged recently.
Survey on publicly available Sinhala Natural Language Processing tools and research
Nisansa de Silva. 2021 · 1906
Earlier work this paper cites.
Animal Farm: A Fairy Story
George Orwell. 1945 · 1945
Earlier work this paper cites.
Assessing endangerment: Expanding Fishman’s GIDS
M Paul Lewis and Gary F Simons. 2010 · 2010
Earlier work this paper cites.
Towards a Sinhala Wordnet
Viraj Welgama, Dulip Lakmal Herath, Chamila Liyanage, Namal Udalamatta, Ruvan Weerasinghe, and Tissa Jayawardana. 2011 · 2011
Earlier work this paper cites.
Automatic speech recognition for under-resourced languages: A survey
Laurent Besacier, Etienne Barnard, Alexey Karpov, and Tanja Schultz. 2014 · 2014
Earlier work this paper cites.
Building a WordNet for Sinhala
Indeewari Wijesiri, Malaka Gallage, Buddhika Gunathilaka, Madhuranga Lakjeewa, Daya Wimalasuriya, Gihan Dias, Rohini Paranavithana, and Nisansa de Silva. 2014 · 2014
Earlier work this paper cites.
Improving website hyperlink structure using server logs
Ashwin Paranjape, Robert West, Leila Zia, and Jure Leskovec. 2016 · 2016
Earlier work this paper cites.
A neural approach to automated essay scoring
Kaveh Taghipour and Hwee Tou Ng. 2016 · 2016
Earlier work this paper cites.
Sindhi language processing: A survey
Wazir Ali Jamro. 2017 · 2017
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M Bender and Batya Friedman. 2018 · 2018
Earlier work this paper cites.
XNLI: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Earlier work this paper cites.
The #Benderrule: On naming the languages we study and why it matters
Emily Bender. 2019 · 2019
Earlier work this paper cites.
The geographic diversity of NLP conferences
Andrew Cains. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
The FLORES evaluation datasets for low-resource machine translation: Nepali–English and Sinhala–English
Francisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, and Marc’Aurelio Ranzato. 2019 · 2019
Earlier work this paper cites.
Keyphrase extraction from disaster-related tweets
Jishnu Ray Chowdhury, Cornelia Caragea, and Doina Caragea. 2019 · 2019
Earlier work this paper cites.
Bootstrapping a neural morphological analyzer for St. Lawrence Island Yupik from a finite-state transducer
Lane Schwartz, Emily Chen, Benjamin Hunt, and Sylvia L.R. Schreiner. 2019 · 2019
Earlier work this paper cites.
Endangered languages meet Modern NLP
Antonios Anastasopoulos, Christopher Cox, Graham Neubig, and Hilaria Cruz. 2020 · 2020
Cited alongside, same era.
Decolonising speech and language technology
Steven Bird. 2020 · 2020
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Cited alongside, same era.
Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020 · 2020
Cited alongside, same era.
The state and fate of linguistic diversity and inclusion in the NLP world
Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020 · 2020
Cited alongside, same era.
Migration of small and endangered languages into the Wikipedia
Armin Hoenen and Marc D Rahn. 2021 · 2021
Later among the works it cites.
When being unseen from mBERT is just the beginning: Handling new languages with multilingual language models
Benjamin Muller, Antonios Anastasopoulos, Benoît Sagot, and Djamé Seddah. 2021 · 2021
Later among the works it cites.
Neural Machine Translation for low-resource languages: A survey
Surangika Ranathunga, En-Shiun Annie Lee, Marjana Prifti Skenduli, Ravi Shekhar, Mehreen Alam, and Rishemjit Kaur. 2021 · 2021
Later among the works it cites.
Fine-tuning self-supervised multilingual sequence-to-sequence models for extremely low-resource NMT
Sarubi Thillainathan, Surangika Ranathunga, and Sanath Jayasena. 2021 · 2021
Later among the works it cites.
Building machine translation systems for the next thousand languages
Ankur Bapna, Isaac Caswell, Julia Kreutzer, Orhan Firat, Daan van Esch, Aditya Siddhant, Mengmeng Niu, Pallavi Baljekar, Xavier Garcia, Wolfgang Macherey, et al. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Divyanshu Kakwani, Anoop Kunchukuttan, Satish Golla, Gokul N.C., Avik Bhattacharyya, Mitesh M. Khapra, and Pratyush Kumar. 2020 · 2020
Cited alongside, same era.
From zero to hero: On the limitations of zero-shot language transfer with multilingual Transformers
Anne Lauscher, Vinit Ravishankar, Ivan Vulić, and Goran Glavaš. 2020 · 2020
Cited alongside, same era.
META-NET white paper series: Key results and cross-language comparison
META-NET. 2020 · 2020
Cited alongside, same era.
UPSTAGE: Unsupervised context augmentation for utterance classification in patient-provider communication
Veronica Perez-Rosas, Shihchen Kuo, William H Herman, Rada Mihalcea, et al. 2020 · 2020
Cited alongside, same era.
Effective approach to develop a sentiment annotator for legal domain in a low resource setting
Gathika Ratnayaka, Nisansa de Silva, Amal Shehan Perera, and Ramesh Pathirana. 2020 · 2020
Cited alongside, same era.
Detecting urgency status of crisis tweets: A transfer learning approach for low resource languages
Efsun Sarioglu Kayi, Linyong Nan, Bohan Qu, Mona Diab, and Kathleen McKeown. 2020 · 2020
Cited alongside, same era.
OPUS-MT – building open translation services for the world
Jörg Tiedemann and Santhosh Thottingal. 2020 · 2020
Cited alongside, same era.
Closest in time.
Local languages, third spaces, and other high-resource scenarios
Steven Bird. 2022 · 2022
Closest in time.
Systematic inequalities in language technology performance across the world’s languages
Damian Blasi, Antonios Anastasopoulos, and Graham Neubig. 2022 · 2022
Closest in time.
‘Reinforces the legitimacy of our language’: Inuktitut officially available on Facebook desktop
News CBC. 2022 · 2022
Closest in time.
BERTifying Sinhala - a comprehensive analysis of pre-trained language models for Sinhala text classification
Vinura Dhananjaya, Piyumal Demotte, Surangika Ranathunga, and Sanath Jayasena. 2022 · 2022
Closest in time.
AmericasNLI: Evaluating zero-shot natural language understanding of pretrained multilingual models in truly low-resource languages
Abteen Ebrahimi, Manuel Mager, Arturo Oncevay, Vishrav Chaudhary, Luis Chiruzzo, Angela Fan, John Ortega, Ricardo Ramos, Annette Rios, Ivan Vladimir Meza Ruiz, Gustavo Giménez-Lugo, Elisabeth Mager, Graham Neubig, Alexis Palmer, Rolando Coto-Solano, Thang Vu, and Katharina Kann. 2022 · 2022
Closest in time.
Dataset geography: Mapping language data to language users
Fahim Faisal, Yinkai Wang, and Antonios Anastasopoulos. 2022 · 2022
Closest in time.
Introducing the digital language equality metric: Contextual factors
Annika Grützner-Zahn and Georg Rehm. 2022 · 2022
Closest in time.
Simran Khanuja, Sebastian Ruder, and Partha Talukdar. 2022 · 2022
Closest in time.
Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, et al. 2022 · 2022
Closest in time.
Adapter based fine-tuning of pre-trained multilingual language models for code-mixed and code-switched text classification
Himashi Rathnayake, Janani Sumanapala, Raveesha Rukshani, and Surangika Ranathunga. 2022 · 2022
Closest in time.
ACL Anthology Corpus with Full Text
Shaurya Rohatgi. 2022 · 2022
Closest in time.