Fetching the paper…
Reading the bibliography…
Natural language processing (NLP) has a significant impact on society via technologies such as machine translation and search engines.
Cross-lingual name tagging and linking for 282 languages
Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji. 2017 · 1958
Earlier work this paper cites.
Ngaju-Dayak Language
Motomitsu Uchibori and Norio Shibata. 1988 · 1988
Earlier work this paper cites.
The bible as a parallel corpus: Annotating the ‘book of 2000 tongues’
Philip Resnik, Mari Broman Olsen, and Mona Diab. 1999 · 2000
Earlier work this paper cites.
Pmindia–a collection of parallel corpora of languages of india
Barry Haddow and Faheem Kirefu. 2020 · 2001
Earlier work this paper cites.
Tata bahasa Jawa mutakhir
Wedhawati Wedhawati, Wiwin E.S.N., Sri Nardiati, Herawati Herawati, Restu Sukesti, Marsono Marsono, Edi Setiyanto, Dirgo Sabariyanto, Syamsul Arifin, Sumadi Sumadi, and Laginem Laginem. 2001 · 2001
Earlier work this paper cites.
Balinese morphosyntax: a lexical-functional approach
I Wayan Arka. 2003 · 2003
Earlier work this paper cites.
The Indonesian language: Its history and role in modern society
James Neil Sneddon. 2003 · 2003
Earlier work this paper cites.
The jrc-acquis: A multilingual aligned parallel corpus with 20+ languages
Ralf Steinberger, Bruno Pouliquen, Anna Widiger, Camelia Ignat, Tomaz Erjavec, Dan Tufis, and Dániel Varga. 2006 · 2006
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, et al. 2007 · 2007
Earlier work this paper cites.
A two-level morphological analyser for the Indonesian language
Femphy Pisceldo, Rahmad Mahendra, Ruli Manurung, and I Wayan Arka. 2008 · 2008
Earlier work this paper cites.
Voice and verb morphology in Minangkabau, a language of West Sumatra, Indonesia
Sophie Elizabeth Crouch. 2009 · 2009
Earlier work this paper cites.
A grammar of Madurese
William D Davies. 2010 · 2010
Earlier work this paper cites.
The case of intercorp, a multilingual parallel corpus
Franti ek Čermák and Alexandr Rosen. 2012 · 2012
Earlier work this paper cites.
Building large monolingual dictionaries at the Leipzig corpora collection: From 100 to 200 languages
Dirk Goldhahn, Thomas Eckart, and Uwe Quasthoff. 2012 · 2012
Earlier work this paper cites.
The Austronesian Languages
Robert Blust et al. 2013 · 2013
Earlier work this paper cites.
Sundanese complementation
Eri Kurniawan. 2013 · 2013
Earlier work this paper cites.
Towards language preservation: Design and collection of graphemically balanced and parallel speech corpora of indonesian ethnic languages
Sakriani Sakti and Satoshi Nakamura. 2013 · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
PanLex: Building a resource for panlingual lexical translation
David Kamholz, Jonathan Pool, and Susan Colowick. 2014 · 2014
Earlier work this paper cites.
Stock price prediction using linear regression based on sentiment analysis
Yahya Eru Cakra and Bayu Distiawan Trisedya. 2015 · 2015
Earlier work this paper cites.
Buzzer detection and sentiment analysis for predicting presidential election results in a twitter nation
Mochamad Ibrahim, Omar Abdillah, Alfan F Wicaksono, and Mirna Adriani. 2015 · 2015
Earlier work this paper cites.
Introduction of the asian language treebank
Hammam Riza, Michael Purwoadi, Teduh Uliniansyah, Aw Ai Ti, Sharifah Mahani Aljunied, Luong Chi Mai, Vu Tat Thang, Nguyen Phuong Thai, Vichet Chea, Sethserey Sam, et al. 2016 · 2016
Earlier work this paper cites.
Lorelei language packs: Data, tools, and resources for technology development in low resource languages
Stephanie Strassel and Jennifer Tracey. 2016 · 2016
Earlier work this paper cites.
Syntactic variation of buginese, a language in austronesian great family
Sukardi Weda. 2016 · 2016
Earlier work this paper cites.
A robust self-learning method for fully unsupervised cross-lingual mappings of word embeddings
Mikel Artetxe, Gorka Labaka, and Eneko Agirre. 2018 · 2018
Earlier work this paper cites.
Prediction and analysis of indonesia presidential election from twitter using sentiment analysis
Widodo Budiharto and Meiliana Meiliana. 2018 · 2018
Cited alongside, same era.
XNLI: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Cited alongside, same era.
Multilingual parallel corpus for global communication plan
Kenji Imamura and Eiichiro Sumita. 2018 · 2018
Cited alongside, same era.
Tufs asian language parallel corpus (talpco)
Hiroki Nomoto, Kenji Okano, David Moeljadi, and Hideo Sawada. 2018 · 2018
Cited alongside, same era.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Cited alongside, same era.
JW300: A wide-coverage parallel corpus for low-resource languages
Željko Agić and Ivan Vulić. 2019 · 2019
Semi-supervised low-resource style transfer of Indonesian informal to formal language with iterative forward-translation
Haryo Akbarianto Wibowo, Tatag Aziz Prawiro, Muhammad Ihsan, Alham Fikri Aji, Radityo Eko Prasojo, Rahmad Mahendra, and Suci Fitriany. 2020 · 2020
Later among the works it cites.
IndoNLU: Benchmark and resources for evaluating Indonesian natural language understanding
Bryan Wilie, Karissa Vincentio, Genta Indra Winata, Samuel Cahyawijaya, Xiaohong Li, Zhi Yuan Lim, Sidik Soleman, Rahmad Mahendra, Pascale Fung, Syafri Bahar, and Ayu Purwarianti. 2020 · 2020
Later among the works it cites.
MasakhaNER: Named Entity Recognition for African Languages
David Ifeoluwa Adelani, Jade Abbott, Graham Neubig, Daniel D’souza, Julia Kreutzer, Constantine Lignos, Chester Palen-Michel, Happy Buzaaba, Shruti Rijhwani, Sebastian Ruder, Stephen Mayhew, Israel Abebe Azime, Shamsuddeen H. Muhammad, Chris Chinenye Emezue, Joyce Nakatumba-Nabende, Perez Ogayo, Aremu Anuoluwapo, Catherine Gitau, Derguene Mbaye, Jesujoba Alabi, Seid Muhie Yimam, Tajuddeen Rabiu Gwadabe, Ignatius Ezeani, Rubungo Andre Niyongabo, Jonathan Mukiibi, Verrah Otiende, Iroro Orife, Davis David, Samba Ngom, Tosin Adewumi, Paul Rayson, Mofetoluwa Adeyemi, Gerald Muriuki, Emmanuel Anebi, Chiamaka Chukwuneke, Nkiruka Odu, Eric Peter Wairagala, Samuel Oyerinde, Clemencia Siro, Tobius Saul Bateesa, Temilola Oloyede, Yvonne Wambui, Victor Akinode, Deborah Nabagereka, Maurice Katusiime, Ayodele Awokoya, Mouhamadane MBOUP, Dibora Gebreyohannes, Henok Tilaye, Kelechi Nwaike, Degaga Wolde, Abdoulaye Faye, Blessing Sibanda, Orevaoghene Ahia, Bonaventure F. P. Dossou, Kelechi Ogueji, Thierno Ibrahima DIOP, Abdoulaye Diallo, Adewale Akinfaderin, Tendai Marengereke, and Salomey Osei. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Multi-label aspect categorization with convolutional neural networks and extreme gradient boosting
Annisa Nurul Azhar, Masayu Leylia Khodra, and Arie Pratama Sutiono. 2019 · 2019
Cited alongside, same era.
Cmu wilderness multilingual speech dataset
Alan W Black. 2019 · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Word2vec model for sentiment analysis of product reviews in indonesian language
M Ali Fauzi. 2019 · 2019
Cited alongside, same era.
The FLORES evaluation datasets for low-resource machine translation: Nepali–English and Sinhala–English
Francisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, and Marc’Aurelio Ranzato. 2019 · 2019
Cited alongside, same era.
Improving bi-lstm performance for indonesian sentiment analysis using paragraph vector
Ayu Purwarianti and Ida Ayu Putu Ari Crisdayanti. 2019 · 2019
Cited alongside, same era.
IndoNLG: Benchmark and resources for evaluating Indonesian natural language generation
Samuel Cahyawijaya, Genta Indra Winata, Bryan Wilie, Karissa Vincentio, Xiaohong Li, Adhiguna Kuncoro, Sebastian Ruder, Zhi Yuan Lim, Syafri Bahar, Masayu Khodra, Ayu Purwarianti, and Pascale Fung. 2021 · 2021
Later among the works it cites.
Ethnologue: Languages of the World. Twenty-fourth edition
David M. Eberhard, Gary F. Simons, and Charles D. Fennig. 2021 · 2021
Later among the works it cites.
BiToD: A bilingual multi-domain dataset for task-oriented dialogue modeling
Zhaojiang Lin, Andrea Madotto, Genta Indra Winata, Peng Xu, Feijun Jiang, Yuxiang Hu, Chen Shi, and Pascale Fung. 2021 · 2021
Later among the works it cites.
IndoNLI: A natural language inference dataset for Indonesian
Rahmad Mahendra, Alham Fikri Aji, Samuel Louvan, Fahrurrozi Rahman, and Clara Vania. 2021 · 2021
Later among the works it cites.
The application of whatsapp to support online learning during the covid-19 pandemic in indonesia
Herri Mulyono, Gunawan Suryoputro, and Shafa Ramadhanya Jamil. 2021 · 2021
Later among the works it cites.
Plan optimization to bilingual dictionary induction for low-resource language families
Arbi Haza Nasution, Yohei Murakami, and Toru Ishida. 2021 · 2021
Later among the works it cites.
Costs to consider in adopting NLP for your business
Made Nindyatama Nityasya, Haryo Akbarianto Wibowo, Radityo Eko Prasojo, and Alham Fikri Aji. 2021 · 2021
Later among the works it cites.
Sentiment analysis on covid19 vaccines in indonesia: From the perspective of sinovac and pfizer
Deden Ade Nurdeni, Indra Budi, and Aris Budi Santoso. 2021 · 2021
Later among the works it cites.
Abusive language and hate speech detection for Indonesian-local language in social media text
Shofianina Dwi Ananda Putri, Muhammad Okky Ibrohim, and Indra Budi. 2021 · 2021
Later among the works it cites.
WikiMatrix: Mining 135M parallel sentences in 1620 language pairs from Wikipedia
Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, and Francisco Guzmán. 2021 · 2021
Later among the works it cites.
IndoCollex: A testbed for morphological transformation of Indonesian word colloquialism
Haryo Akbarianto Wibowo, Made Nindyatama Nityasya, Afra Feyza Akyürek, Suci Fitriany, Alham Fikri Aji, Radityo Eko Prasojo, and Derry Tanti Wijaya. 2021 · 2021
Later among the works it cites.
Language models are few-shot multilingual learners
Genta Indra Winata, Andrea Madotto, Zhaojiang Lin, Rosanne Liu, Jason Yosinski, and Pascale Fung. 2021 · 2021
Later among the works it cites.
mT5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021 · 2021
Later among the works it cites.
One country, 700+ languages: NLP challenges for underrepresented languages and dialects in Indonesia
Alham Fikri Aji, Genta Indra Winata, Fajri Koto, Samuel Cahyawijaya, Ade Romadhony, Rahmad Mahendra, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Timothy Baldwin, Jey Han Lau, and Sebastian Ruder. 2022 · 2022
Closest in time.
AmericasNLI: Evaluating zero-shot natural language understanding of pretrained multilingual models in truly low-resource languages
Abteen Ebrahimi, Manuel Mager, Arturo Oncevay, Vishrav Chaudhary, Luis Chiruzzo, Angela Fan, John Ortega, Ricardo Ramos, Annette Rios, Ivan Vladimir Meza Ruiz, Gustavo Giménez-Lugo, Elisabeth Mager, Graham Neubig, Alexis Palmer, Rolando Coto-Solano, Thang Vu, and Katharina Kann. 2022 · 2022
Closest in time.
NaijaSenti: A nigerian Twitter sentiment corpus for multilingual sentiment analysis
Shamsuddeen Hassan Muhammad, David Adelani, Anuoluwapo Aremu, and Idris Abdulmumin. 2022 · 2022
Closest in time.
Samanantar: The Largest Publicly Available Parallel Corpora Collection for 11 Indic Languages
Gowtham Ramesh, Sumanth Doddapaneni, Aravinth Bheemaraj, Mayank Jobanputra, Raghavan AK, Ajitesh Sharma, Sujit Sahoo, Harshita Diddee, Mahalakshmi J, Divyanshu Kakwani, Navneet Kumar, Aswin Pradeep, Srihari Nagaraj, Kumar Deepak, Vivek Raghavan, Anoop Kunchukuttan, Pratyush Kumar, and Mitesh Shantadevi Khapra. 2022 · 2022
Closest in time.
Writing system and speaker metadata for 2,800+ language varieties
Daan van Esch, Tamar Lucassen, Sebastian Ruder, Isaac Caswell, and Clara E Rivera. 2022 · 2022
Closest in time.
Expanding pretrained models to thousands more languages via lexicon-based adaptation
Xinyi Wang, Sebastian Ruder, and Graham Neubig. 2022 · 2022
Closest in time.
Cross-lingual few-shot learning on unseen languages
Genta Winata, Shijie Wu, Mayank Kulkarni, Thamar Solorio, and Daniel Preoţiuc-Pietro. 2022 · 2022
Closest in time.
Pre-trained transformer-based language models for sundanese
Wilson Wongso, Henry Lucky, and Derwin Suhartono. 2022 · 2022
Closest in time.