Fetching the paper…
Reading the bibliography…
The NLP research community has devoted increased attention to languages beyond English, resulting in considerable improvements for multilingual NLP.
Cross-lingual name tagging and linking for 282 languages
Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji. 2017 · 1958
Earlier work this paper cites.
Some universals of grammar with particular reference to the order of meaningful elements
Joseph Harold Greenberg. 1963 · 1963
Earlier work this paper cites.
A method of language sampling
Jan Rijkhoff, Dik Bakker, Kees Hengeveld, and Peter Kahrel. 1993 · 1993
Earlier work this paper cites.
The comparative method as heuristic
Johanna Nichols. 1996 · 1996
Earlier work this paper cites.
Language sampling
Jan Rijkhoff and Dik Bakker. 1998 · 1998
Earlier work this paper cites.
Co-reference annotation and resources: A multilingual corpus of typologically diverse languages
Felix Sasaki, Claudia Wegener, Andreas Witt, Dieter Metzing, and Jens Pönninghaus. 2002 · 2002
Earlier work this paper cites.
Linguistically naïve != language independent: Why NLP needs linguistic typology
Emily M. Bender. 2009 · 2009
Earlier work this paper cites.
On achieving and evaluating language-independence in NLP
Emily M. Bender. 2011 · 2011
Earlier work this paper cites.
SMALLWorlds – multilingual content-controlled monologues
Peter Juel Henrichsen and Marcus Uneson. 2012 · 2012
Earlier work this paper cites.
Disentangling geography from genealogy
Michael Cysouw. 2013 · 2013
Earlier work this paper cites.
WALS Online (v2020.3)
Matthew S. Dryer and Martin Haspelmath, editors. 2013 · 2013
Earlier work this paper cites.
Capturing diversity in language acquisition research
Sabine Stoll and Balthasar Bickel. 2013 · 2013
Earlier work this paper cites.
A database for measuring linguistic information content
Richard Sproat, Bruno Cartoni, HyunJeong Choe, David Huynh, Linne Ha, Ravindran Rajakumar, and Evelyn Wenzel-Grondie. 2014 · 2014
Earlier work this paper cites.
Sampling for variety
Matti Miestamo, Dik Bakker, and Antti Arppe. 2016 · 2016
Earlier work this paper cites.
URIEL and lang2vec: Representing languages as typological, geographical, and phylogenetic vectors
Patrick Littell, David R. Mortensen, Ke Lin, Katherine Kairis, Carlisle Turner, and Lori Levin. 2017 · 2017
Earlier work this paper cites.
An Introduction to Linguistic Typology , reprinted 2017 with corrections edition
Viveka Velupillai. 2012 · 2017
Earlier work this paper cites.
XNLI: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Earlier work this paper cites.
On the relation between linguistic typology and (limitations of) multilingual language modeling
Daniela Gerz, Ivan Vulić, Edoardo Maria Ponti, Roi Reichart, and Anna Korhonen. 2018 · 2018
Earlier work this paper cites.
Universal dependencies 2.2
Joakim Nivre, Rogier Blokland, Niko Partanen, and Michael Rießler. 2018 · 2018
Earlier work this paper cites.
Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond
Mikel Artetxe and Holger Schwenk. 2019 · 2019
Earlier work this paper cites.
A systematic comparison of methods for low-resource dependency parsing on genuinely low-resource languages
Clara Vania, Yova Kementchedjhieva, Anders Søgaard, and Adam Lopez. 2019 · 2019
Earlier work this paper cites.
On the cross-lingual transferability of monolingual representations
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020 · 2020
Earlier work this paper cites.
Decolonising speech and language technology
Steven Bird. 2020 · 2020
Earlier work this paper cites.
Entity Linking in 100 Languages
Jan A. Botha, Zifei Shan, and Daniel Gillick. 2020 · 2020
Earlier work this paper cites.
Learning and evaluating emotion lexicons for 91 languages
Sven Buechel, Susanna Rücker, and Udo Hahn. 2020 · 2020
Earlier work this paper cites.
TyDi QA: A benchmark for information-seeking question answering in typologically diverse languages
Jonathan H. Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki. 2020 · 2020
Earlier work this paper cites.
Unsupervised cross-lingual part-of-speech tagging for truly low-resource scenarios
Ramy Eskander, Smaranda Muresan, and Michael Collins. 2020b · 2020
Cited alongside, same era.
XHate-999: Analyzing and detecting abusive language across domains and languages
Goran Glavaš, Mladen Karan, and Ivan Vulić. 2020 · 2020
Cited alongside, same era.
The ACQDIV corpus database and aggregation pipeline
Anna Jancso, Steven Moran, and Sabine Stoll. 2020 · 2020
Cited alongside, same era.
X-FACTR: Multilingual factual knowledge retrieval from pretrained language models
Zhengbao Jiang, Antonios Anastasopoulos, Jun Araki, Haibo Ding, and Graham Neubig. 2020 · 2020
Cited alongside, same era.
The state and fate of linguistic diversity and inclusion in the NLP world
Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020 · 2020
Cited alongside, same era.
Weakly Supervised POS Taggers Perform Poorly on Truly Low-Resource Languages
SIGMORPHON–UniMorph 2022 shared task 0: Generalization and typologically diverse morphological inflection
Jordan Kodner, Salam Khalifa, Khuyagbaatar Batsuren, Hossep Dolatian, Ryan Cotterell, Faruk Akkus, Antonios Anastasopoulos, Taras Andrushko, Aryaman Arora, Nona Atanalov, Gábor Bella, Elena Budianskaya, Yustinus Ghanggo Ate, Omer Goldman, David Guriel, Simon Guriel, Silvia Guriel-Agiashvili, Witold Kieraś, Andrew Krizhanovsky, Natalia Krizhanovsky, Igor Marchenko, Magdalena Markowska, Polina Mashkovtseva, Maria Nepomniashchaya, Daria Rodionova, Karina Scheifer, Alexandra Sorova, Anastasia Yemelina, Jeremiah Young, and Ekaterina Vylomova. 2022 · 2022
Later among the works it cites.
Eeny, meeny, miny, moe. how to choose data for morphological inflection
Saliha Muradoglu and Mans Hulden. 2022 · 2022
Later among the works it cites.
xGQA: Cross-lingual visual question answering
Jonas Pfeiffer, Gregor Geigle, Aishwarya Kamath, Jan-Martin Steitz, Stefan Roth, Ivan Vulić, and Iryna Gurevych. 2022 · 2022
Later among the works it cites.
Average is not enough: Caveats of multilingual evaluation
Matúš Pikuliak and Marian Simko. 2022 · 2022
Later among the works it cites.
The “problem” of human label variation: On ground truth in data, modeling and evaluation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Katharina Kann, Ophélie Lacroix, and Anders Søgaard. 2020 · 2020
Cited alongside, same era.
MLQA: Evaluating cross-lingual extractive question answering
Patrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk. 2020 · 2020
Cited alongside, same era.
Manual clustering and spatial arrangement of verbs for multilingual evaluation and typology analysis
Olga Majewska, Ivan Vulić, Diana McCarthy, and Anna Korhonen. 2020 · 2020
Cited alongside, same era.
Morphological segmentation for low resource languages
Justin Mott, Ann Bies, Stephanie Strassel, Jordan Kodner, Caitlin Richter, Hongzhi Xu, and Mitchell Marcus. 2020 · 2020
Cited alongside, same era.
XCOPA: A multilingual dataset for causal commonsense reasoning
Edoardo Maria Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulić, and Anna Korhonen. 2020 · 2020
Cited alongside, same era.
LAReQA: Language-agnostic answer retrieval from a multilingual pool
Uma Roy, Noah Constant, Rami Al-Rfou, Aditya Barua, Aaron Phillips, and Yinfei Yang. 2020 · 2020
Cited alongside, same era.
Multi-SimLex: A large-scale evaluation of multilingual and crosslingual lexical semantic similarity
Ivan Vulić, Simon Baker, Edoardo Maria Ponti, Ulla Petti, Ira Leviant, Kelly Wing, Olga Majewska, Eden Bar, Matt Malone, Thierry Poibeau, Roi Reichart, and Anna Korhonen. 2020 · 2020
Cited alongside, same era.
Barbara Plank. 2022 · 2022
Later among the works it cites.
Square one bias in NLP: Towards a multi-dimensional exploration of the research manifold
Sebastian Ruder, Ivan Vulić, and Anders Søgaard. 2022 · 2022
Later among the works it cites.
DivEMT: Neural machine translation post-editing effort across typologically diverse languages
Gabriele Sarti, Arianna Bisazza, Ana Guerberof-Arenas, and Antonio Toral. 2022 · 2022
Later among the works it cites.
TyDiP: A dataset for politeness classification in nine typologically diverse languages
Anirudh Srinivasan and Eunsol Choi. 2022 · 2022
Later among the works it cites.
Cross-linguistic syntactic difference in multilingual BERT: How good is it and how does it affect transfer?
Ningyu Xu, Tao Gui, Ruotian Ma, Qi Zhang, Jingting Ye, Menghan Zhang, and Xuanjing Huang. 2022 · 2022
Later among the works it cites.
Making more of little data: Improving low-resource automatic speech recognition using data augmentation
Martijn Bartelds, Nay San, Bradley McDonnell, Dan Jurafsky, and Martijn Wieling. 2023 · 2023
Later among the works it cites.
MasakhaPOS: Part-of-speech tagging for typologically diverse African languages
Cheikh M. Bamba Dione, David Ifeoluwa Adelani, Peter Nabende, Jesujoba Alabi, Thapelo Sindane, Happy Buzaaba, Shamsuddeen Hassan Muhammad, Chris Chinenye Emezue, Perez Ogayo, Anuoluwapo Aremu, Catherine Gitau, Derguene Mbaye, Jonathan Mukiibi, Blessing Sibanda, Bonaventure F. P. Dossou, Andiswa Bukula, Rooweither Mabuya, Allahsera Auguste Tapo, Edwin Munkoh-Buabeng, Victoire Memdjokam Koagne, Fatoumata Ouoba Kabore, Amelia Taylor, Godson Kalipe, Tebogo Macucwa, Vukosi Marivate, Tajuddeen Gwadabe, Mboning Tchiaze Elvis, Ikechukwu Onyenwe, Gratien Atindogbe, Tolulope Adelani, Idris Akinade, Olanrewaju Samuel, Marien Nahimana, Théogène Musabeyezu, Emile Niyomutabazi, Ester Chimhenga, Kudzai Gotosa, Patrick Mizha, Apelete Agbolo, Seydou Traore, Chinedu Uchechukwu, Aliyu Yusuf, Muhammad Abdullahi, and Dietrich Klakow. 2023 · 2023
Later among the works it cites.
MASSIVE: A 1M-example multilingual natural language understanding dataset with 51 typologically-diverse languages
Jack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie, Kay Rottmann, Ana Sanchez, Aaron Nash, Liam Urbach, Vishesh Kakarala, Richa Singh, Swetha Ranganath, Laurie Crist, Misha Britan, Wouter Leeuwis, Gokhan Tur, and Prem Natarajan. 2023 · 2023
Later among the works it cites.
Findings of the SIGMORPHON 2023 shared task on interlinear glossing
Michael Ginn, Sarah Moeller, Alexis Palmer, Anna Stacey, Garrett Nicolai, Mans Hulden, and Miikka Silfverberg. 2023 · 2023
Later among the works it cites.
SIGMORPHON–UniMorph 2023 shared task 0: Typologically diverse morphological inflection
Omer Goldman, Khuyagbaatar Batsuren, Salam Khalifa, Aryaman Arora, Garrett Nicolai, Reut Tsarfaty, and Ekaterina Vylomova. 2023 · 2023
Later among the works it cites.
Cross-lingual knowledge distillation for answer sentence selection in low-resource languages
Shivanshu Gupta, Yoshitomo Matsubara, Ankit Chadha, and Alessandro Moschitti. 2023 · 2023
Later among the works it cites.
MultiTACRED: A multilingual version of the TAC relation extraction dataset
Leonhard Hennig, Philippe Thomas, and Sebastian Möller. 2023 · 2023
Later among the works it cites.
Why we need a gradient approach to word order
Natalia Levshina, Savithry Namboodiripad, Marc Allassonnière-Tang, Mathew Kramer, Luigi Talamo, Annemarie Verkerk, Sasha Wilmoth, Gabriela Garrido Rodriguez, Timothy Michael Gupton, Evan Kidd, Zoey Liu, Chiara Naccarato, Rachel Nordlinger, Anastasia Panova, and Natalia Stoynova. 2023 · 2023
Later among the works it cites.
Using neural machine translation for generating diverse challenging exercises for language learner
Frank Palma Gomez, Subhadarshi Panda, Michael Flor, and Alla Rozovskaya. 2023 · 2023
Later among the works it cites.
Understanding compositional data augmentation in typologically diverse morphological inflection
Farhan Samir and Miikka Silfverberg. 2023 · 2023
Later among the works it cites.
Language models are multilingual chain-of-thought reasoners
Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Dipanjan Das, and Jason Wei. 2023 · 2023
Later among the works it cites.
SLABERT talk pretty one day: Modeling second language acquisition with BERT
Aditya Yadavalli, Alekhya Yadavalli, and Vera Tobin. 2023 · 2023
Later among the works it cites.
MIRACL: A multilingual retrieval dataset covering 18 diverse languages
Xinyu Zhang, Nandan Thakur, Odunayo Ogundepo, Ehsan Kamalloo, David Alfonso-Hermelo, Xiaoguang Li, Qun Liu, Mehdi Rezagholizadeh, and Jimmy Lin. 2023 · 2023
Later among the works it cites.
A note on evaluating multilingual benchmarks
Antonis Anastasopoulos. 2019 · 2024
Closest in time.
Multilingual gradient word-order typology from Universal Dependencies
Emi Baylor, Esther Ploeger, and Johannes Bjerva. 2024 · 2024
Closest in time.
A measure for transparent comparison of linguistic diversity in multilingual NLP data sets
Tanja Samardzic, Ximena Gutierrez, Christian Bentz, Steven Moran, and Olga Pelloni. 2024 · 2024
Closest in time.