2020

The State and Fate of Linguistic Diversity and Inclusion in the NLP World

Joshi, Pratik, Santy, Sebastin, Budhiraja, Amar et al.

Understand

Language technologies contribute to promoting multilingualism and linguistic diversity around the world.

  • However, only a very small number of the over 7000 languages of the world are represented in the rapidly evolving language technologies and applications.
  • In this paper we look at the relation between the types of languages, resources, and their representation in NLP conferences to understand the trajectory that different languages have followed over time.
  • Our quantitative investigation underlines the disparity between languages, especially in terms of their resources, and calls into question the "language agnostic" status of current models and systems.

Built on

  • The open language archives community: An infrastructure for distributed archiving of language resources

    Gary Simons and Steven Bird. 2003 · 2003

    Earlier work this paper cites.

  • The ACL anthology reference corpus: A reference dataset for bibliographic research in computational linguistics

    Steven Bird, Robert Dale, Bonnie J Dorr, Bryan Gibson, Mark Thomas Joseph, Min-Yen Kan, Dongwon Lee, Brett Powley, Dragomir R Radev, and Yee Fan Tan. 2008 · 2008

    Earlier work this paper cites.

  • Visualizing data using t-SNE

    Laurens van der Maaten and Geoffrey Hinton. 2008 · 2008

    Earlier work this paper cites.

  • On achieving and evaluating language-independence in NLP

    Emily M. Bender. 2011 · 2011

    Earlier work this paper cites.

  • WALS Online

    Matthew S. Dryer and Martin Haspelmath, editors. 2013 · 2013

    Earlier work this paper cites.

Similar

  • Efficient estimation of word representations in vector space

    Tomas Mikolov, Kai Chen, G. S. Corrado, and J. Dean. 2013 · 2013

    Cited alongside, same era.

  • PanLex: Building a resource for panlingual lexical translation

    David Kamholz, Jonathan Pool, and Susan Colowick. 2014 · 2014

    Cited alongside, same era.

  • SQuAD: 100,000+ questions for machine comprehension of text

    Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016

    Cited alongside, same era.

  • Six challenges for neural machine translation

    Philipp Koehn and Rebecca Knowles. 2017 · 2017

    Cited alongside, same era.

  • Massively multilingual neural machine translation

    Roee Aharoni, Melvin Johnson, and Orhan Firat. 2019 · 2019

    Cited alongside, same era.

Then

  • Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond

    Mikel Artetxe and Holger Schwenk. 2019 · 2019

    Later among the works it cites.

  • Cross-lingual language model pretraining

    Alexis Conneau and Guillaume Lample. 2019 · 2019

    Later among the works it cites.

  • BERT: Pre-training of deep bidirectional transformers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019

    Later among the works it cites.

  • How multilingual is multilingual BERT?

    Telmo Pires, Eva Schlinger, and Dan Garrette. 2019 · 2019

    Later among the works it cites.

  • Modeling language variation and universals: A survey on typological linguistics for natural language processing

    Edoardo Maria Ponti, Helen O’Horan, Yevgeni Berzak, Ivan Vulić, Roi Reichart, Thierry Poibeau, Ekaterina Shutova, and Anna Korhonen. 2019 · 2019

    Later among the works it cites.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…