Fetching the paper…
Reading the bibliography…
Multilingual transformer models like mBERT and XLM-RoBERTa have obtained great improvements for many NLP tasks on a variety of languages.
Cross-lingual name tagging and linking for 282 languages
Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji. 2017 · 1958
Earlier work this paper cites.
Matrix smoothing: A regularization for DNN with transition matrix under noisy labels
Xianbin Lv, Dongxian Wu, and Shu-Tao Xia. 2020 · 2003
Earlier work this paper cites.
Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition
Erik F. Tjong Kim Sang and Fien De Meulder. 2003 · 2003
Earlier work this paper cites.
Anne Lauscher, Vinit Ravishankar, Ivan Vulic, and Goran Glavas. 2020 · 2005
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Çaglar Gülçehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Developing text resources for ten south african languages
Roald Eiselen and Martin J. Puttkammer. 2014 · 2014
Earlier work this paper cites.
Recurrent convolutional neural networks for text classification
Siwei Lai, Liheng Xu, Kang Liu, and Jun Zhao. 2015 · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
Government domain named entity recognition for south African languages
Roald Eiselen. 2016 · 2016
Earlier work this paper cites.
Learning when to trust distant supervision: An application to low-resource POS tagging using cross-lingual projection
Meng Fang and Trevor Cohn. 2016 · 2016
Earlier work this paper cites.
End-to-end sequence labeling via bi-directional LSTM-CNNs-CRF
Xuezhe Ma and Eduard Hovy. 2016 · 2016
Earlier work this paper cites.
LORELEI language packs: Data, tools, and resources for technology development in low resource languages
Stephanie Strassel and Jennifer Tracey. 2016 · 2016
Earlier work this paper cites.
Learning with noise: Enhance distantly supervised relation extraction with dynamic transition matrix
Bingfeng Luo, Yansong Feng, Zheng Wang, Zhanxing Zhu, Songfang Huang, Rui Yan, and Dongyan Zhao. 2017 · 2017
Earlier work this paper cites.
Training a neural network in a low-resource setting on automatically annotated noisy data
Michael A. Hedderich and Dietrich Klakow. 2018 · 2018
Cited alongside, same era.
hauwe: Hausa words embedding for natural language processing
Idris Abdulmumin and Bashir Shehu Galadanci. 2019 · 2019
Cited alongside, same era.
Uncover the ground-truth relations in distant supervision: A neural expectation-maximization framework
Junfan Chen, Richong Zhang, Yongyi Mao, Hongyu Guo, and Jie Xu. 2019 · 2019
Cited alongside, same era.
Targer: Neural argument mining at your fingertips
Artem Chernodub, Oleksiy Oliynyk, Philipp Heidenreich, Alexander Bondarenko, Matthias Hagen, Chris Biemann, and Alexander Panchenko. 2019 · 2019
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Cited alongside, same era.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 2019
Later among the works it cites.
Corpus building for low resource languages in the DARPA LORELEI program
Jennifer Tracey, Stephanie Strassel, Ann Bies, Zhiyi Song, Michael Arrigo, Kira Griffitt, Dana Delgado, Dave Graff, Seth Kulick, Justin Mott, and Neil Kuster. 2019 · 2019
Later among the works it cites.
Learning with noisy labels for sentence-level sentiment classification
Hao Wang, Bing Liu, Chaozhuo Li, Yan Yang, and Tianrui Li. 2019 · 2019
Later among the works it cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew. 2019 · 2019
Later among the works it cites.
Beto, bentz, becas: The surprising cross-lingual effectiveness of BERT
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
mBERT README file
Jacob Devlin. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Ethnologue: Languages of the world. twenty-second edition
David M. Eberhard, Gary F. Simons, and Charles D. Fennig (eds.). 2019 · 2019
Cited alongside, same era.
Towards realistic practices in low-resource natural language processing: The development set
Katharina Kann, Kyunghyun Cho, and Samuel R. Bowman. 2019 · 2019
Cited alongside, same era.
Feature-dependent confusion matrices for low-resource NER labeling with noisy labels
Lukas Lange, Michael A. Hedderich, and Dietrich Klakow. 2019 · 2019
Cited alongside, same era.
Named entity recognition with partially annotated training data
Stephen Mayhew, Snigdha Chaturvedi, Chen-Tse Tsai, and Dan Roth. 2019 · 2019
Cited alongside, same era.
Handling noisy labels for robustly learning from self-training data for low-resource sequence labeling
Debjit Paul, Mittul Singh, Michael A. Hedderich, and Dietrich Klakow. 2019 · 2019
Cited alongside, same era.
Shijie Wu and Mark Dredze. 2019 · 2019
Later among the works it cites.
David Ifeoluwa Adelani, Michael A. Hedderich, Dawei Zhu, Esther van den Berg, and Dietrich Klakow. 2020 · 2020
Closest in time.
Massive vs. Curated Word Embeddings for Low-Resourced Languages. The Case of Yorùbá and Twi
Jesujoba Alabi, Kwabena Amponsah-Kaakyire, David Adelani, and Cristina Espana-Bonet. 2020 · 2020
Closest in time.
Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalization
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020 · 2020
Closest in time.
Weakly supervised pos taggers perform poorly on truly low-resource languages
Katharina Kann, Ophélie Lacroix, and Anders Søgaard. 2020 · 2020
Closest in time.
Universal dependencies v2: An evergrowing multilingual treebank collection
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajič, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, and Daniel Zeman. 2020 · 2020
Closest in time.
Snorkel: rapid training data creation with weak supervision
Alexander Ratner, Stephen H. Bach, Henry R. Ehrenberg, Jason A. Fries, Sen Wu, and Christopher Ré. 2020 · 2020
Closest in time.
Soft gazetteers for low-resource named entity recognition
Shruti Rijhwani, Shuyan Zhou, Graham Neubig, and Jaime Carbonell. 2020 · 2020
Closest in time.