Fetching the paper…
Reading the bibliography…
Language identification (LID) is a crucial precursor for NLP, especially for mining web data.
N-gram-based text categorization
William B. Cavnar and John M. Trenkle. 1994 · 1994
Earlier work this paper cites.
An approach to automatic language identification based on language-dependent phone recognition
Yonghong Yan and E. Barnard. 1995 · 1995
Earlier work this paper cites.
"We just mix": code switching in a South African township
Rosalie Finlayson and Sarah Slabbert. 1997 · 1997
Earlier work this paper cites.
Orthographic diacritics and multilingual computing
John C. Wells. 2000 · 2000
Earlier work this paper cites.
Automatic language identification
Marc A Zissman and Kay M Berkling. 2001 · 2001
Earlier work this paper cites.
African languages and phonological theory
Larry M Hyman. 2003 · 2003
Earlier work this paper cites.
Comparing methods for language identification
Muntsa Padró and Lluís Padró. 2004 · 2004
Earlier work this paper cites.
The Crúbadán project: Corpus building for under-resourced languages
Kevin P. Scannell. 2007 · 2007
Earlier work this paper cites.
Africa as a morphosyntactic area
Denis Creissels, Gerrit J Dimmendaal, Zygmunt Frajzyngier, and Christa König. 2008 · 2008
Earlier work this paper cites.
A comparative study on language identification methods
Lena Grothe, Ernesto William De Luca, and Andreas Nürnberger. 2008 · 2008
Earlier work this paper cites.
Language identification: The long and the short of the matter
Timothy Baldwin and Marco Lui. 2010 · 2010
Earlier work this paper cites.
Accuracy and performance of google’s compact language detector
Michael McCandless. 2010 · 2010
Earlier work this paper cites.
Language detection library for java
Nakatani Shuyo. 2010 · 2010
Earlier work this paper cites.
Cross-domain feature selection for language identification
Marco Lui and Timothy Baldwin. 2011 · 2011
Earlier work this paper cites.
Multilingual sentiment analysis on social media
Erik Tromp. 2011 · 2011
Earlier work this paper cites.
langid.py: An off-the-shelf language identification tool
Marco Lui and Timothy Baldwin. 2012 · 2012
Earlier work this paper cites.
Robust language identification in short, noisy texts: Improvements to liga
John Vogel and David Tresner-Kirsch. 2012 · 2012
Earlier work this paper cites.
Selecting and weighting n-grams to identify 1100 languages
Ralf D. Brown. 2013 · 2013
Earlier work this paper cites.
WALS Online
Matthew S. Dryer and Martin Haspelmath, editors. 2013 · 2013
Earlier work this paper cites.
Query word labeling and back transliteration for Indian languages: Shared task system description
Spandana Gella, Jatin Sharma, and Kalika Bali. 2013 · 2013
Earlier work this paper cites.
Learning phrase representations using RNN encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
“ye word kis lang ka hai bhai?” testing the limits of word level language identification
Spandana Gella, Kalika Bali, and Monojit Choudhury. 2014 · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014 · 2014
Earlier work this paper cites.
Merging comparable data sources for the discrimination of similar languages: The dsl corpus collection
Liling Tan, Marcos Zampieri, Nikola Ljubešic, and Jörg Tiedemann. 2014 · 2014
Earlier work this paper cites.
A report on the DSL shared task 2014
Marcos Zampieri, Liling Tan, Nikola Ljubešić, and Jörg Tiedemann. 2014 · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyung Hyun Cho, and Yoshua Bengio. 2015 · 2015
Cited alongside, same era.
What is this? is it code switching, code mixing or language alternating?
D R Mabule. 2015 · 2015
Cited alongside, same era.
Overview of the DSL shared task 2015
Marcos Zampieri, Liling Tan, Nikola Ljubešić, Jörg Tiedemann, and Preslav Nakov. 2015 · 2015
Cited alongside, same era.
Simple tools for exploring variation in code-switching for linguists
Gualberto A. Guzman, Jacqueline Serigos, Barbara E. Bullock, and Almeida Jacqueline Toribio. 2016 · 2016
Cited alongside, same era.
Language identification for South African Bantu languages using rank order statistics
Meluleki Dube and Hussein Suleman. 2019 · 2019
Later among the works it cites.
ParaCrawl: Web-scale parallel corpora for the languages of the EU
Miquel Esplà, Mikel Forcada, Gema Ramírez-Sánchez, and Hieu Hoang. 2019 · 2019
Later among the works it cites.
Automatic language identification in texts: A survey
Tommi Jauhiainen, Marco Lui, Marcos Zampieri, Timothy Baldwin, and Krister Lindén. 2019 · 2019
Later among the works it cites.
Asynchronous pipelines for processing huge corpora on medium to low resource infrastructures
Pedro Javier Ortiz Suárez, Benoît Sagot, and Laurent Romary. 2019 · 2019
Later among the works it cites.
Toward micro-dialect identification in diaglossic and code-switched environments
Muhammad Abdul-Mageed, Chiyu Zhang, AbdelRahim Elmadany, and Lyle Ungar. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Armand Joulin, Edouard Grave, Piotr Bojanowski, Matthijs Douze, Hérve Jégou, and Tomas Mikolov. 2016 · 2016
Cited alongside, same era.
Discriminating between similar languages and Arabic dialect identification: A report on the third DSL shared task
Shervin Malmasi, Marcos Zampieri, Nikola Ljubešić, Preslav Nakov, Ahmed Ali, and Jörg Tiedemann. 2016 · 2016
Cited alongside, same era.
Identification of languages in Algerian Arabic multilingual documents
Wafia Adouane and Simon Dobnik. 2017 · 2017
Cited alongside, same era.
Improving the character ngram model for the DSL task with BM25 weighting and less frequently used feature sets
Yves Bestgen. 2017 · 2017
Cited alongside, same era.
A dataset and classifier for recognizing social media English
Su Lin Blodgett, Johnny Wei, and Brendan O’Connor. 2017 · 2017
Cited alongside, same era.
Analysis and prediction of Dutch-English code-switching in Dutch social media messages
N. Dongen. 2017 · 2017
Cited alongside, same era.
Improved text language identification for the South African languages
Bernardt Duvenhage, Mfundo Ntini, and Phala Ramonyai. 2017a · 2017
Cited alongside, same era.
Massive vs. curated embeddings for low-resourced languages: the case of Yorùbá and Twi
Jesujoba Alabi, Kwabena Amponsah-Kaakyire, David Adelani, and Cristina España-Bonet. 2020 · 2020
Later among the works it cites.
Language ID in the wild: Unexpected challenges on the path to a thousand-language web text corpus
Isaac Caswell, Theresa Breiner, Daan van Esch, and Ankur Bapna. 2020 · 2020
Later among the works it cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Later among the works it cites.
Mapping languages: the corpus of global language use
Jonathan Dunn. 2020 · 2020
Later among the works it cites.
CCAligned: A massive collection of cross-lingual web-document pairs
Ahmed El-Kishky, Vishrav Chaudhary, Francisco Guzmán, and Philipp Koehn. 2020 · 2020
Later among the works it cites.
A monolingual approach to contextualized word embeddings for mid-resource languages
Pedro Javier Ortiz Suárez, Laurent Romary, and Benoît Sagot. 2020 · 2020
Later among the works it cites.
Exploring Amharic sentiment analysis from social media texts: Building annotation tools and classification models
Seid Muhie Yimam, Hizkiel Mitiku Alemayehu, Abinew Ayele, and Chris Biemann. 2020 · 2020
Later among the works it cites.
ARBERT & MARBERT: Deep bidirectional transformers for Arabic
Muhammad Abdul-Mageed, AbdelRahim Elmadany, and El Moatez Billah Nagoudi. 2021 · 2021
Later among the works it cites.
Ethnologue: Languages of the world
David M Eberhard, F Simons Gary, and Charles D Fennig (eds). 2021 · 2021
Later among the works it cites.
Quality at a glance: An audit of web-crawled multilingual datasets
Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suárez, Iroro Orife, Kelechi Ogueji, Andre Niyongabo Rubungo, Toan Q. Nguyen, Mathias Müller, André Müller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Kudugunta, Yacine Jernite, Mathias Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine Çabuk Ballı, Stella Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, and Mofetoluwa Adeyemi. 2021 · 2021
Later among the works it cites.
WikiMatrix: Mining 135M parallel sentences in 1620 language pairs from Wikipedia
Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, and Francisco Guzmán. 2021 · 2021
Later among the works it cites.
Transformer based language identification for malayalam-english code-mixed text
S. Thara and Prabaharan Poornachandran. 2021 · 2021
Later among the works it cites.
Improved language identification through cross-lingual self-supervised learning
Andros Tjandra, Diptanu Gon Choudhury, Frank Zhang, Kritika Singh, Alexis Conneau, Alexei Baevski, Assaf Sela, Yatharth Saraf, and Michael Auli. 2021 · 2021
Later among the works it cites.
mT5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021 · 2021
Later among the works it cites.
Towards afrocentric NLP for African languages: Where we are and where we can go
Ife Adebara and Muhammad Abdul-Mageed. 2022 · 2022
Closest in time.
Ancestor-to-creole transfer is not a walk in the park
Heather Lent, Emanuele Bugliarello, and Anders Søgaard. 2022 · 2022
Closest in time.
Naijasenti: A nigerian twitter sentiment corpus for multilingual sentiment analysis
Shamsuddeen Hassan Muhammad, David Ifeoluwa Adelani, Sebastian Ruder, Ibrahim Said Ahmad, Idris Abdulmumin, Bello Shehu Bello, Monojit Choudhury, Chris Chinenye Emezue, Saheed Salahudeen Abdullahi, Anuoluwapo Aremu, Alipio Jeorge, and Pavel Brazdil. 2022 · 2022
Closest in time.
AraT5: Text-to-text transformers for Arabic language generation
El Moatez Billah Nagoudi, AbdelRahim Elmadany, and Muhammad Abdul-Mageed. 2022 · 2022
Closest in time.