Fetching the paper…
Reading the bibliography…
Although Indonesian is known to be the fourth most frequently used language over the internet, the research progress on this language in the natural language processing (NLP) is slow-moving due to a lack of available resources.
Cross-lingual language model pretraining
Guillaume Lample and Alexis Conneau. 2019 · 1901
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1911
Earlier work this paper cites.
Clue: A chinese language understanding evaluation benchmark
Liang Xu, Xuanwei Zhang, Lu Li, Hai Hu, Chenjie Cao, Weitang Liu, Junyi Li, Yudong Li, Kai Sun, Yechen Xu, et al. 2020 · 2004
Earlier work this paper cites.
A machine learning approach for indonesian question answering system
Ayu Purwarianti, Masatoshi Tsuchiya, and Seiichi Nakagawa. 2007 · 2007
Earlier work this paper cites.
A two-level morphological analyser for the indonesian language
Femphy Pisceldo, Rahmad Mahendra, Ruli Manurung, and I Wayan Arka. 2008 · 2008
Earlier work this paper cites.
Usage of indonesian possessive verbal predicates: a statistical analysis based on questionnaire and storytelling surveys
David Moeljadi. 2012 · 2012
Earlier work this paper cites.
Designing an indonesian part of speech tagset and manually tagged indonesian corpus
Arawinda Dinakaramani, Fam Rashel, Andry Luthfi, and Ruli Manurung. 2014 · 2014
Earlier work this paper cites.
Fasttext.zip: Compressing text classification models
Armand Joulin, Edouard Grave, Piotr Bojanowski, Matthijs Douze, Hérve Jégou, and Tomas Mikolov. 2016 · 2016
Earlier work this paper cites.
Opensubtitles2016: Extracting large parallel corpora from movie and tv subtitles
Pierre Lison and Jörg Tiedemann. 2016 · 2016
Earlier work this paper cites.
CoNLL 2017 shared task - automatically annotated raw texts and word embeddings
Filip Ginter, Jan Hajič, Juhani Luotolahti, Milan Straka, and Daniel Zeman. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Learning word vectors for 157 languages
Edouard Grave, Piotr Bojanowski, Prakhar Gupta, Armand Joulin, and Tomas Mikolov. 2018 · 2018
Cited alongside, same era.
Investigating bi-lstm and crf with pos tag embedding for indonesian named entity tagger
Devin Hoesen and Ayu Purwarianti. 2018 · 2018
Cited alongside, same era.
Aspect detection and sentiment classification using deep neural network for indonesian aspect-based sentiment analysis
Arfinda Ilmania, Samuel Cahyawijaya, Ayu Purwarianti, et al. 2018 · 2018
Cited alongside, same era.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Cited alongside, same era.
Toward a standardized and more accurate indonesian part-of-speech tagging
Kemal Kurniawan and Alham Fikri Aji. 2018 · 2018
Cited alongside, same era.
Pre-training with whole word masking for chinese bert
Yiming Cui, Wanxiang Che, Ting Liu, Bing Qin, Ziqing Yang, Shijin Wang, and Guoping Hu. 2019 · 2019
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Jordhy Fernando, Masayu Leylia Khodra, and Ali Akbar Septiandri. 2019 · 2019
Later among the works it cites.
To lemmatize or not to lemmatize: How word normalisation affects ELMo performance in word sense disambiguation
Andrey Kutuzov and Elizaveta Kuzmenko. 2019 · 2019
Later among the works it cites.
Flaubert: Unsupervised language model pre-training for french
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hiroki Nomoto, Kenji Okano, David Moeljadi, and Hideo Sawada. 2018 · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Emotion classification on indonesian twitter dataset
Mei Silviana Saputri, Rahmad Mahendra, and Mirna Adriani. 2018 · 2018
Cited alongside, same era.
Semi-supervised textual entailment on indonesian wikipedia data
Ken Nabila Setya and Rahmad Mahendra. 2018 · 2018
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
Jw300: A wide-coverage parallel corpus for low-resource languages
Željko Agić and Ivan Vulić. 2019 · 2019
Cited alongside, same era.
Multi-label aspect categorization with convolutional neural networks and extreme gradient boosting
A. N. Azhar, M. L. Khodra, and A. P. Sutiono. 2019 · 2019
Cited alongside, same era.
Hang Le, Loïc Vial, Jibril Frej, Vincent Segonne, Maximin Coavoux, Benjamin Lecouteux, Alexandre Allauzen, Benoît Crabbé, Laurent Besacier, and Didier Schwab. 2019 · 2019
Later among the works it cites.
Improving joint layer rnn based keyphrase extraction by using syntactical features
Miftahul Mahfuzh, Sidik Soleman, and Ayu Purwarianti. 2019 · 2019
Later among the works it cites.
Camembert: a tasty french language model
Louis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont, Laurent Romary, Éric Villemonte de la Clergerie, Djamé Seddah, and Benoît Sagot. 2019 · 2019
Later among the works it cites.
Asynchronous Pipeline for Processing Huge Corpora on Medium to Low Resource Infrastructures
Pedro Javier Ortiz Suárez, Benoît Sagot, and Laurent Romary. 2019 · 2019
Later among the works it cites.
Improving bi-lstm performance for indonesian sentiment analysis using paragraph vector
Ayu Purwarianti and Ida Ayu Putu Ari Crisdayanti. 2019 · 2019
Later among the works it cites.
Ali Akbar Septiandri and Arie Pratama Sutiono. 2019 · 2019
Later among the works it cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019 · 2019
Later among the works it cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020 · 2020
Closest in time.