Fetching the paper…
Reading the bibliography…
State-of-the-art natural language processing (NLP) models are trained on massive training corpora, and report a superlative performance on evaluation datasets.
Language Family Relationship Preserved in Non-native English. In COLING . Dublin City University and Association for Computational Linguistics, 1940–1949
Ryo Nagata. 2014 · 1949
Earlier work this paper cites.
Dialect, language, nation 1
Einar Haugen. 1966 · 1966
Earlier work this paper cites.
Interactions of Written and Spoken Hindi
RS Yadav. 1974 · 1974
Earlier work this paper cites.
Discourse strategies
John J Gumperz. 1982 · 1982
Earlier work this paper cites.
Toward a theory of social dialect variation
Anthony S Kroch. 1986 · 1986
Earlier work this paper cites.
Dialectometry: a short overview of the principles and practice of quantitative classification of linguistic atlas data. In Contributions to Quantitative Linguistics: First International Conference on Quantitative Linguistics, QUALICO, Trier, 1991 . Springer, 277–315
Hans Goebl. 1993 · 1991
Earlier work this paper cites.
The other tongue: English across cultures
Braj B Kachru. 1992 · 1992
Earlier work this paper cites.
Measuring dialect distance phonetically. In Computational phonology: third meeting of the acl special interest group in computational phonology
John Nerbonne and Wilbert Heeringa. 1997 · 1997
Earlier work this paper cites.
The Vocabulary of Australian English
Bruce Moore. 1999 · 1999
Earlier work this paper cites.
Computational comparison and classification of dialects
John Nerbonne and Wilbert Heeringa. 2001 · 2001
Earlier work this paper cites.
Computational Comparison and Classification of Dialects
J Nerbonne and Wilbert Heeringa. 2002 · 2002
Earlier work this paper cites.
Parsing arabic dialects. In 11th Conference of the European Chapter of the Association for Computational Linguistics . 369–376
David Chiang, Mona Diab, Nizar Habash, Owen Rambow, and Safiullah Shareef. 2006 · 2006
Earlier work this paper cites.
Dialect Classification for Online Podcasts Fusing Acoustic and Language Based Structural and Semantic Information. In ACL . 21–24
Rahul Chitturi and John Hansen. 2008 · 2006
Earlier work this paper cites.
The acoustic characteristics of/hVd/vowels in the speech of some Australian teenagers
Felicity Cox. 2006 · 2006
Earlier work this paper cites.
MAGEAD: A Morphological Analyzer and Generator for the Arabic Dialects. In ACL . 681–688
Nizar Habash and Owen Rambow. 2006 · 2006
Earlier work this paper cites.
Australian English
Felicity Cox and Sallyanne Palethorpe. 2007 · 2007
Earlier work this paper cites.
A Layered Grammar Model: Using Tree-Adjoining Grammars to Build a Common Syntactic Kernel for Related Dialects. In Ninth International Workshop on Tree Adjoining Grammar and Related Frameworks (TAG+9) . Tübingen, Germany, 157–164
Pascal Vaillant. 2008 · 2008
Earlier work this paper cites.
English as a lingua franca: Interpretations and attitudes
Jennifer Jenkins. 2009 · 2009
Earlier work this paper cites.
Incorporating Dialectal Variability for Socially Equitable Language Identification. In ACL . Vancouver, Canada, 51–57
David Jurgens, Yulia Tsvetkov, and Dan Jurafsky. 2017 · 2009
Earlier work this paper cites.
COLABA: Arabic dialect annotation and processing. In Workshop on semitic language processing . 66–74
Mona Diab, Nizar Habash, Owen Rambow, Mohamed Altantawy, and Yassine Benajiba. 2010 · 2010
Earlier work this paper cites.
Example-based Translation of Japanese Functional Expressions utilizing Semantic Equivalence Classes. In 4th Workshop on Patent Translation . Xiamen, China
Yusuke Abe, Takafumi Suzuki, Bing Liang, Takehito Utsuro, Mikio Yamamoto, Suguru Matsuyoshi, and Yasuhide Kawada. 2011 · 2011
Earlier work this paper cites.
Dialect translation: integrating Bayesian co-segmentation models with pivot-based SMT. In Workshop on Algorithms and Resources for Modelling of Dialects and Language Varieties
Michael Paul, Andrew Finch, Paul Dixon, and Eiichiro Sumita. 2011 · 2011
Earlier work this paper cites.
Simplified guidelines for the creation of Large Scale Dialectal Arabic Annotations.. In LREC . 371–378
Heba Elfardy and Mona T Diab. 2012 · 2012
Earlier work this paper cites.
Im/politeness across Englishes
Michael Haugh and Klaus P Schneider. 2012 · 2012
Earlier work this paper cites.
langid.py: An Off-the-shelf Language Identification Tool. In ACL System Demonstrations . Jeju Island, Korea, 25–30
Marco Lui and Timothy Baldwin. 2012 · 2012
Earlier work this paper cites.
Getting stuff done: Comparing e-mail requests from students in higher education in Britain and Australia
Andrew John Merrison, Jack J Wilson, Bethan L Davies, and Michael Haugh. 2012 · 2012
Earlier work this paper cites.
Appropriate behaviour across varieties of English
Klaus P Schneider. 2012 · 2012
Earlier work this paper cites.
Machine Translation of Arabic Dialects. In NAACL . Montréal, Canada, 49–59
Rabih Zbib, Erika Malchiodi, Jacob Devlin, David Stallard, Spyros Matsoukas, Richard Schwartz, John Makhoul, Omar F. Zaidan, and Chris Callison-Burch. 2012 · 2012
Earlier work this paper cites.
Building bilingual lexicon to create Dialect Tunisian corpora and adapt language model. In Second Workshop on Hybrid Approaches to Translation . Sofia, Bulgaria, 88–93
Rahma Boujelbane, Mariem Ellouze khemekhem, Siwar BenAyed, and Lamia Hadrich Belguith. 2013 · 2013
Earlier work this paper cites.
Classifying English documents by national dialect. In Australasian Language Technology Association Workshop . 5–15
Marco Lui and Paul Cook. 2013 · 2013
Earlier work this paper cites.
Sana: A large scale multi-genre, multi-dialect lexicon for arabic subjectivity and sentiment analysis.. In LREC . 1162–1169
Muhammad Abdul-Mageed and Mona T Diab. 2014 · 2014
Earlier work this paper cites.
Unsupervised Word Segmentation Improves Dialectal Arabic to English Machine Translation. In Workshop on Arabic Natural Language Processing . ACL, Doha, Qatar, 207–216
Kamla Al-Mannai, Hassan Sajjad, Alaa Khader, Fahad Al Obaidli, Preslav Nakov, and Stephan Vogel. 2014 · 2014
Earlier work this paper cites.
A multidialectal parallel corpus of Arabic. In LREC 2014 . European Language Resources Association (ELRA), 1240–1245
Houda Bouamor, Nizar Habash, and Kemal Oflazer. 2014 · 2014
Earlier work this paper cites.
A Multi-Dialect, Multi-Genre Corpus of Informal Written Arabic.. In LREC . 241–245
Ryan Cotterell and Chris Callison-Burch. 2014 · 2014
Earlier work this paper cites.
Verifiably Effective Arabic Dialect Identification. In EMNLP . 1465–1468
Kareem Darwish, Hassan Sajjad, and Hamdy Mubarak. 2014 · 2014
Earlier work this paper cites.
Predicting Dialect Variation in Immigrant Contexts Using Light Verb Constructions. In EMNLP . 1391–1395
A. Seza Doğruöz and Preslav Nakov. 2014 · 2014
Earlier work this paper cites.
AusTalk: an audio-visual corpus of Australian English. In Ninth International Conference on Language Resources and Evaluation (LREC’14) . European Language Resources Association (ELRA), 3105–3109
Dominique Estival, Steve Cassidy, Felicity Cox, and Denis Burnham. 2014 · 2014
Earlier work this paper cites.
Domain and Dialect Adaptation for Machine Translation into Egyptian Arabic. In Workshop on Arabic Natural Language Processing . Association for Computational Linguistics, Doha, Qatar, 196–206
Serena Jeblee, Weston Feely, Houda Bouamor, Alon Lavie, Nizar Habash, and Kemal Oflazer. 2014 · 2014
Earlier work this paper cites.
Developing an Egyptian Arabic Treebank: Impact of Dialectal Morphology on Annotation and Tool Development.. In LREC . 2348–2354
Mohamed Maamouri, Ann Bies, Seth Kulick, Michael Ciul, Nizar Habash, and Ramy Eskander. 2014 · 2014
Earlier work this paper cites.
The culture map: Breaking through the invisible boundaries of global business
Erin Meyer. 2014 · 2014
Earlier work this paper cites.
Sentence Level Dialect Identification for Machine Translation System Selection. In 52nd ACL (Volume 2: Short Papers) . Baltimore, Maryland, 772–778
Wael Salloum, Heba Elfardy, Linda Alamir-Salloum, Nizar Habash, and Mona Diab. 2014 · 2014
Earlier work this paper cites.
A report on the DSL shared task 2014. In first workshop on applying NLP tools to similar languages, varieties and dialects . 58–67
Marcos Zampieri, Liling Tan, Nikola Ljubešić, and Jörg Tiedemann. 2014 · 2014
Earlier work this paper cites.
Challenges of studying and processing dialects in social media. In Workshop on Noisy User-generated Text . 9–18
Anna Jørgensen, Dirk Hovy, and Anders Søgaard. 2015 · 2015
Earlier work this paper cites.
Hedonic or utilitarian? Exploring the impact of communication style alignment on user’s perception of virtual health advisory services
Manning Li and Jiye Mao. 2015 · 2015
Earlier work this paper cites.
Machine Translation Experiments on PADIC: A Parallel Arabic DIalect Corpus. In PACLIC . Shanghai, China, 26–34
Karima Meftouh, Salima Harrat, Salma Jamoussi, Mourad Abbas, and Kamel Smaili. 2015 · 2015
Earlier work this paper cites.
Dialects
Todd L Sandel. 2015 · 2015
Earlier work this paper cites.
Natural language processing for dialectical Arabic: A survey. In Arabic Natural Language Processing Workshop . 36–48
Abdulhadi Shoufan and Sumaya Alameri. 2015 · 2015
Earlier work this paper cites.
Building Monolingual Word Alignment Corpus for the Greater China Region. In Workshop on Language Technology for Closely Related Languages, Varieties and Dialects . Association for Computational Linguistics, Hissar, Bulgaria, 85–94
Fan Xu, Xiongfei Xu, Mingwen Wang, and Maoxi Li. 2015 · 2015
Earlier work this paper cites.
Overview of the DSL shared task 2015. In Workshop on Language Technology for Closely Related Languages, Varieties and Dialects . 1–9
Marcos Zampieri, Liling Tan, Nikola Ljubešić, Jörg Tiedemann, and Preslav Nakov. 2015 · 2015
Earlier work this paper cites.
Botta: An arabic dialect chatbot. In COLING (System Demonstrations) . 208–212
Dana Abu Ali and Nizar Habash. 2016 · 2016
Earlier work this paper cites.
Demographic Dialectal Variation in Social Media: A Case Study of African-American English. In EMNLP . Austin, Texas, 1119–1130
Su Lin Blodgett, Lisa Green, and Brendan O’Connor. 2016 · 2016
Earlier work this paper cites.
Creating resources for Dialectal Arabic from a single annotation: A case study on Egyptian and Levantine. In COLING . 3455–3465
Ramy Eskander, Nizar Habash, Owen Rambow, and Arfath Pasha. 2016 · 2016
Earlier work this paper cites.
Discriminating similar languages: Evaluations and explorations. In LREC
Cyril Goutte, Serge Léger, Shervin Malmasi, and Marcos Zampieri. 2016 · 2016
Earlier work this paper cites.
A large scale corpus of Gulf Arabic. In LREC . European Language Resources Association (ELRA), 4282–4289
Salam Khalifa, Nizar Habash, Dana Abdulrahim, and Sara Hassan. 2016 · 2016
Earlier work this paper cites.
Discriminating between similar languages and arabic dialect identification: A report on the third dsl shared task. In VarDial . 1–14
Shervin Malmasi, Marcos Zampieri, Nikola Ljubešić, Preslav Nakov, Ahmed Ali, and Jörg Tiedemann. 2016 · 2016
Earlier work this paper cites.
Translating Dialectal Arabic as Low Resource Language using Word Embedding. In RANLP . INCOMA Ltd., Varna, Bulgaria, 52–57
Ebtesam H Almansor and Ahmed Al-Ani. 2017 · 2017
Earlier work this paper cites.
Alg/fr: A step by step construction of a lexicon between algerian dialect and french. In PACLIC , Vol. 31
Faical Azouaou and Imane Guellil. 2017 · 2017
Earlier work this paper cites.
A Morphological Parser for Odawa. In Workshop on the Use of Computational Methods in the Study of Endangered Languages . 1–9
Dustin Bowers, Antti Arppe, Jordan Lachler, Sjur Moshagen, and Trond Trosterud. 2017 · 2017
Earlier work this paper cites.
Discriminating between similar languages with word-level convolutional neural networks. In VarDial . 124–130
Marcelo Criscuolo and Sandra Aluisio. 2017 · 2017
Earlier work this paper cites.
Low Resourced Machine Translation via Morpho-syntactic Modeling: The Case of Dialectal Arabic. In Machine Translation Summit XVI: Research Track . Nagoya Japan, 185–200
Alexander Erdmann, Nizar Habash, Dima Taji, and Houda Bouamor. 2017 · 2017
Earlier work this paper cites.
Synthetic Data for Neural Machine Translation of Spoken-Dialects. In 14th International Conference on Spoken Language Translation . International Workshop on Spoken Language Translation, Tokyo, Japan, 82–89
Hany Hassan, Mostafa Elaraby, and Ahmed Y. Tawfik. 2017 · 2017
Earlier work this paper cites.
Kurdish Interdialect Machine Translation. In Fourth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial) . ACL, Valencia, Spain, 63–72
Hossein Hassani. 2017 · 2017
Earlier work this paper cites.
Curras: an annotated corpus for the Palestinian Arabic dialect
Mustafa Jarrar, Nizar Habash, Faeq Alrimawi, Diyam Akra, and Nasser Zalmout. 2017 · 2017
Earlier work this paper cites.
Reframing AI discourse
Deborah G Johnson and Mario Verdicchio. 2017 · 2017
Earlier work this paper cites.
Sentiment analysis of tunisian dialects: Linguistic ressources and experiments. In Arabic Natural Language Processing Workshop . 55–61
Salima Mdhaffar, Fethi Bougares, Yannick Esteve, and Lamia Hadrich-Belguith. 2017 · 2017
Earlier work this paper cites.
Identifying the Authors’ National Variety of English in Social Media Texts. In International Conference Recent Advances in Natural Language Processing, RANLP 2017 . INCOMA Ltd., Varna, Bulgaria, 671–678
Vasiliki Simaki, Panagiotis Simakis, Carita Paradis, and Andreas Kerren. 2017 · 2017
Earlier work this paper cites.
Universal Dependencies Parsing for Colloquial Singaporean English. In ACL . 1732–1744
Hongmin Wang, Yue Zhang, GuangYong Leonard Chan, Jie Yang, and Hai Leong Chieu. 2017 · 2017
Earlier work this paper cites.
You tweet what you speak: A city-level dataset of arabic dialects. In LREC
Muhammad Abdul-Mageed, Hassan Alhuzali, and Mohamed Elaraby. 2018 · 2018
Earlier work this paper cites.
Multi-dialect Neural Machine Translation and Dialectometry. In PACLIC . ACL, Hong Kong
Kaori Abe, Yuichiroh Matsubayashi, Naoaki Okazaki, and Kentaro Inui. 2018 · 2018
Earlier work this paper cites.
Creating an Arabic Dialect Text Corpus by Exploring Twitter, Facebook, and Online Newspapers. In OSACT 3: The 3rd Workshop on Open-Source Arabic Corpora and Processing Tools . 54
Areej Alshutayri and Eric Atwell. 2018 · 2018
Earlier work this paper cites.
Towards enhancement of a lexicon-based approach for Saudi dialect sentiment analysis
Adel Assiri, Ahmed Emam, and Hmood Al-Dossari. 2018 · 2018
Earlier work this paper cites.
Twitter Universal Dependency Parsing for African-American and Mainstream American English. In ACL . 1415–1425
Su Lin Blodgett, Johnny Wei, and Brendan O’Connor. 2018 · 2018
Earlier work this paper cites.
The MADAR arabic dialect corpus and lexicon. In LREC
Houda Bouamor, Nizar Habash, Mohammad Salameh, Wajdi Zaghouani, Owen Rambow, Dana Abdulrahim, Ossama Obeid, Salam Khalifa, Fadhl Eryani, Alexander Erdmann, et al · 2018
Earlier work this paper cites.
Multi-dialect Arabic POS tagging: a CRF approach. In LREC . European Language Resources Association (ELRA), 93–98
Kareem Darwish, Hamdy Mubarak, Mohamed Eldesouki, Ahmed Abdelali, Younes Samih, Randah Alharbi, Mohammed Attia, Walid Magdy, and Laura Kallmeyer. 2018 · 2018
Earlier work this paper cites.
Arsas: An arabic speech-act and sentiment corpus of tweets
A Elmadany, Hamdy Mubarak, and Walid Magdy. 2018b · 2018
Earlier work this paper cites.
Addressing noise in multidialectal word embeddings. In ACL . 558–565
Alexander Erdmann, Nasser Zalmout, and Nizar Habash. 2018 · 2018
Cited alongside, same era.
Maghrebi Arabic dialect processing: an overview
Salima Harrat, Karima Meftouh, and Kamel Smaïli. 2018 · 2018
Cited alongside, same era.
Machine Translation of Low-Resource Spoken Dialects: Strategies for Normalizing Swiss German. In LREC . European Language Resources Association (ELRA), Miyazaki, Japan
Pierre-Edouard Honnet, Andrei Popescu-Belis, Claudiu Musat, and Michael Baeriswyl. 2018 · 2018
Cited alongside, same era.
Capturing Regional Variation with Distributed Place Representations and Geographic Retrofitting. In EMNLP . ACL, Brussels, Belgium, 4383–4394
Dirk Hovy and Christoph Purschke. 2018 · 2018
Cited alongside, same era.
MADARi: A Web Interface for Joint Arabic Morphological Annotation and Spelling Correction. In LREC . European Language Resources Association (ELRA), Miyazaki, Japan
Ossama Obeid, Salam Khalifa, Nizar Habash, Houda Bouamor, Wajdi Zaghouani, and Kemal Oflazer. 2018 · 2018
A comprehensive review of arabic text summarization
Asmaa Elsaid, Ammar Mohammed, Lamiaa Fattouh Ibrahim, and Mohammed M Sakre. 2022 · 2022
Later among the works it cites.
The Effect of Arabic Dialect Familiarity on Data Annotation. In Arabic Natural Language Processing Workshop . 399–408
Ibrahim Abu Farha and Walid Magdy. 2022 · 2022
Later among the works it cites.
AraConv: Developing an Arabic task-oriented dialogue system using multi-lingual transformer model mT5
Ahlam Fuad and Maha Al-Yahya. 2022 · 2022
Later among the works it cites.
ArzEn-ST: A Three-way Speech Translation Corpus for Code-Switched Egyptian Arabic-English. In Seventh Arabic Natural Language Processing Workshop (WANLP) . Association for Computational Linguistics, Abu Dhabi, United Arab Emirates (Hybrid), 119–130
Injy Hamed, Nizar Habash, Slim Abdennadher, and Ngoc Thang Vu. 2022 · 2022
Later among the works it cites.
Exploring the role of grammar and word choice in bias toward african american english (aae) in hate speech classification. In 2022 ACM Conference on Fairness, Accountability, and Transparency . 789–798
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Fine-grained Arabic dialect identification. In COLING
Mohammad Salameh, Houda Bouamor, and Nizar Habash. 2018 · 2018
Cited alongside, same era.
Arsentd-lev: A multi-topic corpus for target-based sentiment analysis in arabic levantine tweets
Ramy Baly, Alaa Khaddaj, Hazem Hajj, Wassim El-Hajj, and Khaled Bashir Shaban. 2019 · 2019
Cited alongside, same era.
Modeling Global Syntactic Variation in English Using Dialect Classification. In VarDial . Association for Computational Linguistics, 42–53
Jonathan Dunn. 2019 · 2019
Cited alongside, same era.
OlloBot-towards a text-based arabic health conversational agent: Evaluation and results. In RANLP . 295–303
Ahmed Fadhil et al · 2019
Cited alongside, same era.
Automatic language identification in texts: A Survey
Tommi Jauhiainen, Marco Lui, Marcos Zampieri, Timothy Baldwin, and Krister Lindén. 2019 · 2019
Cited alongside, same era.
Arabic dialogue act recognition for textual chatbot systems. In First International Workshop on NLP Solutions for Under Resourced Languages . 43–49
Alaa Joukhadar, Huda Saghergy, Leen Kweider, and Nada Ghneim. 2019 · 2019
Cited alongside, same era.
Syntax-ignorant N-gram embeddings for sentiment analysis of Arabic dialects. In Arabic Natural Language Processing Workshop . 30–39
Hala Mulki, Hatem Haddad, Mourad Gridach, and Ismail Babaoğlu. 2019 · 2019
Cited alongside, same era.
Camille Harris, Matan Halevy, Ayanna Howard, Amy Bruckman, and Diyi Yang. 2022 · 2022
Later among the works it cites.
A weak supervised transfer learning approach for sentiment analysis to the Kuwaiti dialect. In The Seventh Arabic Natural Language Processing Workshop (WANLP) . 161–173
Fatemah Husain, Hana Al-Ostad, and Halima Omar. 2022 · 2022
Later among the works it cites.
Morphosyntactic Tagging with Pre-trained Language Models for Arabic and its Dialects. In Findings of ACL . 1708–1719
Go Inoue, Salam Khalifa, and Nizar Habash. 2022 · 2022
Later among the works it cites.
Early Guessing for Dialect Identification. In Findings of EMNLP . 6417–6426
Vani Kanjirangat, Tanja Samardzic, Fabio Rinaldi, and Ljiljana Dolamic. 2022 · 2022
Later among the works it cites.
SAIDS: A Novel Approach for Sentiment Analysis Informed of Dialect and Sarcasm
Abdelrahman Kaseb and Mona Farouk. 2022 · 2022
Later among the works it cites.
The Norwegian Dialect Corpus Treebank. In LREC . 4827–4832
Andre Kåsen, Kristin Hagen, Anders Nøklestad, Joel Priestly, Per Erik Solberg, and Dag Trygve Truslew Haug. 2022 · 2022
Later among the works it cites.
Singlish Message Paraphrasing: A Joint Task of Creole Translation and Text Normalization. In COLING . 3924–3936
Zhengyuan Liu, Shikang Ni, Ai Ti Aw, and Nancy F. Chen. 2022 · 2022
Later among the works it cites.
Exploring Methods for Building Dialects-Mandarin Code-Mixing Corpora: A Case Study in Taiwanese Hokkien. In Findings of EMNLP . 6287–6305
Sin-En Lu, Bo-Han Lu, Chao-Yi Lu, and Richard Tzong-Han Tsai. 2022 · 2022
Later among the works it cites.
AAEBERT: Debiasing BERT-based Hate Speech Detection Models via Adversarial Learning. In ICMLA . IEEE, 1606–1612
Ebuka Okpala, Long Cheng, Nicodemus Mbwambo, and Feng Luo. 2022 · 2022
Later among the works it cites.
Analyzing the Dialect Diversity in Multi-document Summaries. In COLING . 6208–6221
Olubusayo Olabisi, Aaron Hudson, Antonie Jetter, and Ameeta Agrawal. 2022 · 2022
Later among the works it cites.
Dealing with dialects in literary translation: Problems and strategies
Al-Khanji Rajai and Narjes Ennasser. 2022 · 2022
Later among the works it cites.
A Semi-supervised Approach for a Better Translation of Sentiment in Dialectical Arabic UGT. In Arabic Natural Language Processing Workshop . Association for Computational Linguistics, Abu Dhabi, United Arab Emirates (Hybrid)
Hadeel Saadany, Constantin Orăsan, Emad Mohamed, and Ashraf Tantawy. 2022 · 2022
Later among the works it cites.
Unsupervised Arabic dialect segmentation for machine translation
Wael Salloum and Nizar Habash. 2022 · 2022
Later among the works it cites.
Perceptual Overlap in Classification of L2 Vowels: Australian English Vowels Perceived by Experienced Mandarin Listeners. In PACLIC . 317–324
Yizhou Wang, Rikke L. Bundgaard-Nielsen, Brett J. Baker, and Olga Maxwell. 2022 · 2022
Later among the works it cites.
VALUE: Understanding Dialect Disparity in NLU. In ACL . Dublin, Ireland, 3701–3720
Caleb Ziems, Jiaao Chen, Camille Harris, Jessica Anderson, and Diyi Yang. 2022 · 2022
Later among the works it cites.
NADI 2023: The Fourth Nuanced Arabic Dialect Identification Shared Task. In ArabicNLP . Association for Computational Linguistics, Singapore (Hybrid), 600–613
Muhammad Abdul-Mageed, AbdelRahim Elmadany, Chiyu Zhang, El Moatez Billah Nagoudi, Houda Bouamor, and Nizar Habash. 2023 · 2023
Later among the works it cites.
Findings of the VarDial Evaluation Campaign 2023. In VarDial . Association for Computational Linguistics, Dubrovnik, Croatia, 251–261
Noëmi Aepli, Çağrı Çöltekin, Rob Van Der Goot, Tommi Jauhiainen, Mourhaf Kazzaz, Nikola Ljubešić, Kai North, Barbara Plank, Yves Scherrer, and Marcos Zampieri. 2023 · 2023
Later among the works it cites.
Do All Languages Cost the Same? Tokenization in the Era of Commercial Language Models. In EMNLP . 9904–9923
Orevaoghene Ahia, Sachin Kumar, Hila Gonen, Jungo Kasai, David R Mortensen, Noah A Smith, and Yulia Tsvetkov. 2023 · 2023
Later among the works it cites.
Benchmarking Dialectal Arabic-Turkish Machine Translation. In Machine Translation Summit XIX . Asia-Pacific Association for Machine Translation, Macau SAR, China, 261–271
Hasan Alkheder, Houda Bouamor, Nizar Habash, and Ahmet Zengin. 2023 · 2023
Later among the works it cites.
Low-resource Bilingual Dialect Lexicon Induction with Large Language Models. In 24th Nordic Conference on Computational Linguistics (NoDaLiDa) . University of Tartu Library, Tórshavn, Faroe Islands, 371–385
Katya Artemova and Barbara Plank. 2023 · 2023
Later among the works it cites.
Cross-lingual Strategies for Low-resource Language Modeling: A Study on Five Indic Dialects. In Actes de CORIA-TALN 2023. Actes de la 30e Conférence sur le Traitement Automatique des Langues Naturelles (TALN), volume 1 : travaux de recherche originaux – articles longs . Paris, France, 28–42
Niyati Bafna, Cristina España-Bonet, Josef Van Genabith, Benoît Sagot, and Rachel Bawden. 2023 · 2023
Later among the works it cites.
Does Manipulating Tokenization Aid Cross-Lingual Transfer? A Study on POS Tagging for Non-Standardized Languages. In VarDial . 40–54
Verena Blaschke, Hinrich Schütze, and Barbara Plank. 2023 · 2023
Later among the works it cites.
Double modals in contemporary British and Irish speech
Steven Coats. 2023 · 2023
Later among the works it cites.
The MT@ BZ Corpus: machine translation & legal language. In Annual Conference of the European Association for Machine Translation . 171–180
Flavia De Camillis, Egon Waldemar Stemle, Elena Chiocchetti, and Francesco Fernicola. 2023 · 2023
Later among the works it cites.
MultiSpider: towards benchmarking multilingual text-to-SQL semantic parsing. In AAAI , Vol. 37. 12745–12753
Longxu Dou, Yan Gao, Mingyang Pan, Dingzirui Wang, Wanxiang Che, Dechen Zhan, and Jian-Guang Lou. 2023 · 2023
Later among the works it cites.
MD3: The Multi-Dialect Dataset of Dialogues. In Proc. INTERSPEECH 2023 . 4059–4063
Jacob Eisenstein, Vinodkumar Prabhakaran, Clara Rivera, Dorottya Demszky, and Devyani Sharma. 2023 · 2023
Later among the works it cites.
TADA : Task Agnostic Dialect Adapters for English. In Findings of ACL . 813–824
William Held, Caleb Ziems, and Diyi Yang. 2023 · 2023
Later among the works it cites.
Quantifying the Dialect Gap and its Correlates Across Languages. In Findings of EMNLP . 7226–7245
Anjali Kantharuban, Ivan Vulić, and Anna Korhonen. 2023 · 2023
Later among the works it cites.
ALDi: Quantifying the Arabic Level of Dialectness of Text. In EMNLP . 10597–10611
Amr Keleg, Sharon Goldwater, and Walid Magdy. 2023 · 2023
Later among the works it cites.
Murreviikko - A Dialectologically Annotated and Normalized Dataset of Finnish Tweets. In VarDial . Association for Computational Linguistics, Dubrovnik, Croatia, 31–39
Olli Kuparinen. 2023 · 2023
Later among the works it cites.
Dialect-to-Standard Normalization: A Large-Scale Multilingual Evaluation. In Findings of EMNLP . Association for Computational Linguistics, 13814–13828
Olli Kuparinen, Aleksandra Miletić, and Yves Scherrer. 2023 · 2023
Later among the works it cites.
A Measure for Linguistic Coherence in Spatial Language Variation. In VarDIAL . 133–141
Alfred Lameli and Andreas Schönberg. 2023 · 2023
Later among the works it cites.
A Parallel Corpus for Vietnamese Central-Northern Dialect Text Transfer. In Findings of EMNLP . 13839–13855
Thang Le and Anh Luu. 2023 · 2023
Later among the works it cites.
DADA: Dialect Adaptation via Dynamic Aggregation of Linguistic Rules. In EMNLP . Singapore, 13776–13793
Yanchen Liu, William Held, and Diyi Yang. 2023 · 2023
Later among the works it cites.
Kaushal Kumar Maurya, Rahul Kejriwal, Maunendra Sankar Desarkar, and Anoop Kunchukuttan. 2023 · 2023
Later among the works it cites.
NormMark: A Weakly Supervised Markov Model for Socio-cultural Norm Discovery. In Findings of ACL . Toronto, Canada, 5081–5089
Farhad Moghimifar, Shilin Qu, Tongtong Wu, Yuan-Fang Li, and Gholamreza Haffari. 2023 · 2023
Later among the works it cites.
Double modals in Australian and New Zealand English
Cameron Morin and Steven Coats. 2023 · 2023
Later among the works it cites.
A comprehensive overview of large language models
Humza Naveed, Asad Ullah Khan, Shi Qiu, Muhammad Saqib, Saeed Anwar, Muhammad Usman, Naveed Akhtar, Nick Barnes, and Ajmal Mian. 2023 · 2023
Later among the works it cites.
STT4SG-350: A Speech Corpus for All Swiss German Dialect Regions. In ACL (Short Papers) . Toronto, Canada, 1763–1772
Michel Plüss, Jan Deriu, Yanick Schraner, Claudio Paonessa, Julia Hartmann, Larissa Schmidt, Christian Scheller, Manuela Hürlimann, Tanja Samardžić, Manfred Vogel, and Mark Cieliebak. 2023 · 2023
Later among the works it cites.
GeoLingIt at EVALITA 2023: Overview of the Geolocation of Linguistic Variation in Italy Task. In Evaluation Campaign of Natural Language Processing and Speech Tools for Italian . CEUR.org, Parma, Italy
Alan Ramponi and Camilla Casula. 2023b · 2023
Later among the works it cites.
Enriching the NArabizi Treebank: A Multifaceted Approach to Supporting an Under-Resourced Language. In 17th Linguistic Annotation Workshop (LAW-XVII) . Association for Computational Linguistics, Toronto, Canada, 266–278
Arij Riabi, Menel Mahamdi, and Djamé Seddah. 2023 · 2023
Later among the works it cites.
FRMT: A Benchmark for Few-Shot Region-Aware Machine Translation
Parker Riley, Timothy Dozat, Jan A. Botha, Xavier Garcia, Dan Garrette, Jason Riesa, Orhan Firat, and Noah Constant. 2023 · 2023
Later among the works it cites.
Dialect-robust Evaluation of Generated Text. In ACL . Toronto, Canada, 6010–6028
Jiao Sun, Thibault Sellam, Elizabeth Clark, Tu Vu, Timothy Dozat, Dan Garrette, Aditya Siddhant, Jacob Eisenstein, and Sebastian Gehrmann. 2023 · 2023
Later among the works it cites.
Task-Agnostic Low-Rank Adapters for Unseen English Dialects. In Findings of ACL
Zedian Xiao, William Held, Yanchen Liu, and Diyi Yang. 2023 · 2023
Later among the works it cites.
Socialdial: A benchmark for socially-aware dialogue systems. In ACM SIGIR . 2712–2722
Haolan Zhan, Zhuang Li, Yufei Wang, Linhao Luo, Tao Feng, Xiaoxi Kang, Yuncheng Hua, Lizhen Qu, Lay-Ki Soon, Suraj Sharma, et al · 2023
Later among the works it cites.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al · 2023
Later among the works it cites.
Multi-VALUE: A Framework for Cross-Dialectal English NLP. In ACL . 744–768
Caleb Ziems, William Held, Jingfeng Yang, Jwala Dhamala, Rahul Gupta, and Diyi Yang. 2023 · 2023
Later among the works it cites.
Voices Unheard: NLP Resources and Models for Yorùbá Regional Dialects. In EMNLP . Association for Computational Linguistics, Miami, Florida, USA, 4392–4409
Orevaoghene Ahia, Anuoluwapo Aremu, Diana Abagyan, Hila Gonen, David Ifeoluwa Adelani, Daud Abolade, Noah A. Smith, and Yulia Tsvetkov. 2024 · 2024
Closest in time.
CODET: A Benchmark for Contrastive Dialectal Evaluation of Machine Translation. In Findings of EACL . 1790–1859
Md Mahfuz Ibn Alam, Sina Ahmadi, and Antonios Anastasopoulos. 2024 · 2024
Closest in time.
Exploring the Robustness of Task-oriented Dialogue Systems for Colloquial German Varieties. In EACL) . 445–468
Ekaterina Artemova, Verena Blaschke, and Barbara Plank. 2024 · 2024
Closest in time.
What do dialect speakers want? a survey of attitudes towards language technology for german dialects
Verena Blaschke, Christoph Purschke, Hinrich Schütze, and Barbara Plank. 2024 · 2024
Closest in time.
VarDial Evaluation Campaign 2024: Commonsense Reasoning in Dialects and Multi-Label Similar Language Identification. In VarDial . Association for Computational Linguistics, Mexico City, Mexico, 1–15
Adrian-Gabriel Chifu, Goran Glavaš, Radu Tudor Ionescu, Nikola Ljubešić, Aleksandra Miletić, Filip Miletić, Yves Scherrer, and Ivan Vulić. 2024 · 2024
Closest in time.
Machine Translation Of Marathi Dialects: A Case Study Of Kadodi. In Eleventh Workshop on Asian Translation (WAT 2024) . Association for Computational Linguistics, Miami, Florida, USA, 36–44
Raj Dabre, Mary Dabre, and Teresa Pereira. 2024 · 2024
Closest in time.
Multi-Dialect Vietnamese: Task, Dataset, Baseline Models and Challenges. In EMNLP . Association for Computational Linguistics, Miami, Florida, USA, 7476–7498
Nguyen Van Dinh, Thanh Chi Dang, Luan Thanh Nguyen, and Kiet Van Nguyen. 2024 · 2024
Closest in time.
DIALECTBENCH: A NLP Benchmark for Dialects, Varieties, and Closely-Related Languages. In ACL
Fahim Faisal, Orevaoghene Ahia, Aarohi Srivastava, Kabir Ahuja, David Chiang, Yulia Tsvetkov, and Antonios Anastasopoulos. 2024 · 2024
Closest in time.
Linguistic Bias in ChatGPT: Language Models Reinforce Dialect Discrimination. In EMNLP . Association for Computational Linguistics, Miami, Florida, USA, 13541–13564
Eve Fleisig, Genevieve Smith, Madeline Bossi, Ishita Rustagi, Xavier Yin, and Dan Klein. 2024 · 2024
Closest in time.
Modeling Gender and Dialect Bias in Automatic Speech Recognition. In Findings of EMNLP . Association for Computational Linguistics, Miami, Florida, USA, 15166–15184
Camille Harris, Chijioke Mgbahurike, Neha Kumar, and Diyi Yang. 2024 · 2024
Closest in time.
Dialect prejudice predicts AI decisions about people’s character, employability, and criminality
Valentin Hofmann, Pratyusha Ria Kalluri, Dan Jurafsky, and Sharese King. 2024 · 2024
Closest in time.
Enhancing Neural Machine Translation for Ainu-Japanese: A Comprehensive Study on the Impact of Domain and Dialect Integration. In 4th International Conference on Natural Language Processing for Digital Humanities . Association for Computational Linguistics, Miami, USA, 413–422
Ryo Igarashi and So Miyagawa. 2024 · 2024
Closest in time.
CreoleVal: Multilingual Multitask Benchmarks for Creoles
Heather Lent, Kushal Tatariya, Raj Dabre, Yiyi Chen, Marcell Fekete, Esther Ploeger, Li Zhou, Ruth-Ann Armstrong, Abee Eijansantos, Catriona Malau, Hans Erik Heje, Ernests Lavrinovics, Diptesh Kanojia, Paul Belony, Marcel Bollmann, Loïc Grobol, Miryam de Lhoneux, Daniel Hershcovich, Michel DeGraff, Anders Søgaard, and Johannes Bjerva. 2024 · 2024
Closest in time.
Examining Language Modeling Assumptions Using an Annotated Literary Dialect Corpus. In 4th International Conference on Natural Language Processing for Digital Humanities . Association for Computational Linguistics, Miami, USA, 325–330
Craig Messner and Thomas Lippincott. 2024 · 2024
Closest in time.
The Zeno’s Paradox of ‘Low-Resource’ Languages. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, Miami, Florida, USA, 17753–17774
Hellina Hailu Nigatu, Atnafu Lambebo Tonja, Benjamin Rosman, Thamar Solorio, and Monojit Choudhury. 2024 · 2024
Closest in time.
What do Large Language Models Need for Machine Translation Evaluation?. In Proceedings of the 2024 Conference on EMNLP . Association for Computational Linguistics, Miami, Florida, USA, 3660–3674
Shenbin Qian, Archchana Sindhujan, Minnie Kabra, Diptesh Kanojia, Constantin Orasan, Tharindu Ranasinghe, and Fred Blain. 2024 · 2024
Closest in time.
Language Varieties of Italy: Technology Challenges and Opportunities
Alan Ramponi. 2024 · 2024
Closest in time.
Casablanca: Data and Models for Multidialectal Arabic Speech Recognition. In EMNLP . Association for Computational Linguistics, Miami, Florida, USA, 21745–21758
Bashar Talafha, Karima Kadaoui, Samar Mohamed Magdy, Mariem Habiboullah, Chafei Mohamed Chafei, Ahmed Oumar El-Shangiti, Hiba Zayed, Mohamedou Cheikh Tourad, Rahaf Alhamouri, Rwaa Assi, Aisha Alraeesi, Hour Mohamed, Fakhraddin Alwajih, Abdelrahman Mohamed, Abdellah El Mekki, El Moatez Billah Nagoudi, Benelhadj Djelloul Mama Saadia, Hamzah A. Alsayadi, Walid Al-Dhabyani, Sara Shatnawi, Yasir Ech-chammakhy, Amal Makouar, Yousra Berrachedi, Mustafa Jarrar, Shady Shehata, Ismail Berrada, and Muhammad Abdul-Mageed. 2024 · 2024
Closest in time.
Cross-Dialectal Transfer and Zero-Shot Learning for Armenian Varieties: A Comparative Analysis of RNNs, Transformers and LLMs. In 4th International Conference on Natural Language Processing for Digital Humanities . Association for Computational Linguistics, Miami, USA, 438–449
Chahan Vidal-Gorène, Nadi Tomeh, and Victoria Khurshudyan. 2024 · 2024
Closest in time.
Extracting Lexical Features from Dialects via Interpretable Dialect Classifiers. In NAACL . 54–69
Roy Xie, Orevaoghene Ahia, Yulia Tsvetkov, and Antonios Anastasopoulos. 2024 · 2024
Closest in time.
Findings of the Quality Estimation Shared Task at WMT 2024: Are LLMs Closing the Gap in QE?. In Proceedings of the Ninth WMT . ACL, Miami, Florida, USA, 82–109
Chrysoula Zerva, Frederic Blain, José G. C. De Souza, Diptesh Kanojia, Sourabh Deoghare, Nuno M. Guerreiro, Giuseppe Attanasio, Ricardo Rei, Constantin Orasan, Matteo Negri, Marco Turchi, Rajen Chatterjee, Pushpak Bhattacharyya, Markus Freitag, and André Martins. 2024 · 2024
Closest in time.
RENOVI: A Benchmark Towards Remediating Norm Violations in Socio-Cultural Conversations
Haolan Zhan, Zhuang Li, Xiaoxi Kang, Tao Feng, Yuncheng Hua, Lizhen Qu, Yi Ying, Mei Rianto Chandra, Kelly Rosalin, Jureynolds Jureynolds, et al · 2024
Closest in time.
Creating a Lexicon of Bavarian Dialect by Means of Facebook Language Data and Crowdsourcing. In LREC . 2029–2033
Manuel Burghardt, Daniel Granvogl, and Christian Wolff. 2016 · 2033
Closest in time.