Fetching the paper…
Reading the bibliography…
Lexical normalization, the translation of non-canonical data to standard language, has shown to improve the performance of manynatural language processing tasks on social media.
Discourse strategies , volume 1
John J Gumperz. 1982 · 1982
Earlier work this paper cites.
Social motivations for codeswitching: Evidence from Africa
Carol Myers-Scotton. 1995 · 1995
Earlier work this paper cites.
Random forests
Leo Breiman. 2001 · 2001
Earlier work this paper cites.
A systematic comparison of various statistical alignment models
Franz Josef Och and Hermann Ney. 2003 · 2003
Earlier work this paper cites.
Massive Choice, Ample tasks (MaChAmp): A toolkit for multi-task learning in NLP
Rob van der Goot, Ahmet Üstün, Alan Ramponi, Ibrahim Sharaf, and Barbara Plank. 2021 · 2005
Earlier work this paper cites.
A phrase-based statistical model for SMS text normalization
AiTi Aw, Min Zhang, Juan Xiao, and Jian Su. 2006 · 2006
Earlier work this paper cites.
Tweetmotif: Exploratory search and topic summarization for twitter
Brendan O’Connor, Michel Krieger, and David Ahn. 2010 · 2010
Earlier work this paper cites.
Lexical normalisation of short text messages: Makn sens a #twitter
Bo Han and Timothy Baldwin. 2011 · 2011
Earlier work this paper cites.
A character-level machine translation approach for normalization of SMS abbreviations
Deana Pennell and Yang Liu. 2011 · 2011
Earlier work this paper cites.
The Cambridge handbook of linguistic code-switching
Almeida Jacqueline Toribio and Barbara E Bullock. 2012 · 2012
Earlier work this paper cites.
Polyglot: Distributed word representations for multilingual nlp
Rami Al-Rfou, Bryan Perozzi, and Steven Skiena. 2013 · 2013
Earlier work this paper cites.
Introducción a la tarea compartida Tweet-Norm 2013: Normalización léxica de tuits en español
Inaki Alegria, Nora Aranberri, Víctor Fresno, Pablo Gamallo, Lluis Padró, Inaki San Vicente, Jordi Turmo, and Arkaitz Zubiaga. 2013 · 2013
Earlier work this paper cites.
Twitter part-of-speech tagging for all: Overcoming sparse and noisy data
Leon Derczynski, Alan Ritter, Sam Clark, and Kalina Bontcheva. 2013 · 2013
Earlier work this paper cites.
What to do about bad language on the internet
Jacob Eisenstein. 2013 · 2013
Earlier work this paper cites.
Universal dependency annotation for multilingual parsing
Ryan McDonald, Joakim Nivre, Yvonne Quirmbach-Brundage, Yoav Goldberg, Dipanjan Das, Kuzman Ganchev, Keith Hall, Slav Petrov, Hao Zhang, Oscar Täckström, Claudia Bedini, Núria Bertomeu Castelló, and Jungmee Lee. 2013 · 2013
Earlier work this paper cites.
Efficient higher-order CRFs for morphological tagging
Thomas Mueller, Helmut Schmid, and Hinrich Schütze. 2013 · 2013
Earlier work this paper cites.
Adaptive parser-centric text normalization
Congle Zhang, Tyler Baldwin, Howard Ho, Benny Kimelfeld, and Yunyao Li. 2013 · 2013
Cited alongside, same era.
Identifying languages at the word level in code-mixed Indian social media text
Amitava Das and Björn Gambäck. 2014 · 2014
Cited alongside, same era.
Improving the utility of social media with Natural Language Processing
Bo Han. 2014 · 2014
Cited alongside, same era.
Shared tasks of the 2015 workshop on noisy user-generated text: Twitter lexical normalization and named entity recognition
Timothy Baldwin, Marie Catherine de Marneffe, Bo Han, Young-Bum Kim, Alan Ritter, and Wei Xu. 2015 · 2015
Cited alongside, same era.
NCSU-SAS-Ning: Candidate generation and feature engineering for supervised lexical normalization
Ning Jin. 2015 · 2015
Cited alongside, same era.
Overview of FIRE-2015 shared task on mixed script information retrieval
Parser adaptation for social media by integrating normalization
Rob van der Goot and Gertjan van Noord. 2017 · 2017
Later among the works it cites.
To normalize, or not to normalize: The impact of normalization on part-of-speech tagging
Rob van der Goot, Barbara Plank, and Malvina Nissim. 2017 · 2017
Later among the works it cites.
Bag of tricks for efficient text classification
Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. 2017 · 2017
Later among the works it cites.
Lexical normalization of roman Urdu text
Zareen Sharf and Saif Ur Rahman. 2017 · 2017
Later among the works it cites.
Universal Dependency parsing for Hindi-English code-switching
Irshad Bhat, Riyaz A. Bhat, Manish Shrivastava, and Dipti Sharma. 2018 · 2018
Later among the works it cites.
Word translation without parallel data
Guillaume Lample, Alexis Conneau, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Royal Sequiera, Monojit Choudhury, Parth Gupta, Paolo Rosso, Shubham Kumar, Somnath Banerjee, Sudip Kumar Naskar, Sivaji Bandyopadhyay, Gokul Chittaranjan, Amitava Das, et al. 2015 · 2015
Cited alongside, same era.
Part of speech tagging for code switched data
Fahad AlGhamdi, Giovanni Molina, Mona Diab, Thamar Solorio, Abdelati Hawwari, Victor Soto, and Julia Hirschberg. 2016 · 2016
Cited alongside, same era.
A Turkish-German code-switching corpus
Özlem Çetinoğlu. 2016 · 2016
Cited alongside, same era.
Part of speech annotation of a Turkish-German code-switching corpus
Özlem Çetinoğlu and Çağrı Çöltekin. 2016 · 2016
Cited alongside, same era.
Normalising Slovene data: historical texts vs. user-generated content
Nikola Ljubešic, Katja Zupan, Darja Fišer, and Tomaz Erjavec. 2016 · 2016
Cited alongside, same era.
Overview for the second shared task on language identification in code-switched data
Giovanni Molina, Fahad AlGhamdi, Mahmoud Ghoneim, Abdelati Hawwari, Nicolas Rey-Villamizar, Mona Diab, and Thamar Solorio. 2016 · 2016
Cited alongside, same era.
Universal dependencies v1: A multilingual treebank collection
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Yoav Goldberg, Jan Hajič, Christopher D. Manning, Ryan McDonald, Slav Petrov, Sampo Pyysalo, Natalia Silveira, Reut Tsarfaty, and Daniel Zeman. 2016 · 2016
Cited alongside, same era.
Later among the works it cites.
Joint part-of-speech and language ID tagging for code-switched data
Victor Soto and Julia Hirschberg. 2018 · 2018
Later among the works it cites.
A fast, compact, accurate model for language identification of codemixed text
Yuan Zhang, Jason Riesa, Daniel Gillick, Anton Bakalov, Jason Baldridge, and David Weiss. 2018 · 2018
Later among the works it cites.
Normalising non-standardised orthography in Algerian code-switched user-generated data
Wafia Adouane, Jean-Philippe Bernardy, and Simon Dobnik. 2019 · 2019
Later among the works it cites.
Normalization of Indonesian-English code-mixed twitter data
Anab Maulana Barik, Rahmad Mahendra, and Mirna Adriani. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
MoNoise: A multi-lingual and easy-to-use lexical normalization tool
Rob van der Goot. 2019 · 2019
Later among the works it cites.
Adapting sequence to sequence models for text normalization in social media
Ismini Lourentzou, Kabir Manghnani, and ChengXiang Zhai. 2019 · 2019
Later among the works it cites.
Enhancing BERT for lexical normalization
Benjamin Muller, Benoit Sagot, and Djamé Seddah. 2019 · 2019
Later among the works it cites.
Norm it! lexical normalization for Italian and its downstream effects for dependency parsing
Rob van der Goot, Alan Ramponi, Tommaso Caselli, Cafagna Michele, and Lorenzo De Mattei. 2020 · 2020
Closest in time.
Universal Dependencies v2: An evergrowing multilingual treebank collection
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajič, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, and Daniel Zeman. 2020 · 2020
Closest in time.