Fetching the paper…
Reading the bibliography…
While natural language processing tools have been developed extensively for some of the world's languages, a significant portion of the world's over 7000 languages are still neglected.
1901
Earlier work this paper cites.
1907
Earlier work this paper cites.
1909
Earlier work this paper cites.
Pan X, Zhang B, May J, et al (2017) Cross-lingual name tagging and linking for 282 languages. In: Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, Vancouver, Canada, pp 1946–1958, 10.18653/v1/P17-1178 , URL https://aclanthology.org/P17-1178
1958
Earlier work this paper cites.
Clark K, Luong MT, Le QV, et al (2020) Electra: Pre-training text encoders as discriminators rather than generators. 2003.10555
2003
Earlier work this paper cites.
2003
Earlier work this paper cites.
Lewis DD, Yang Y, Russell-Rose T, et al (2004) Rcv1: A new benchmark collection for text categorization research. Journal of machine learning research 5(Apr):361–397
2004
Earlier work this paper cites.
Koehn P (2005) Europarl: A parallel corpus for statistical machine translation. In: Proceedings of Machine Translation Summit X: Papers, Phuket, Thailand, pp 79–86, URL https://aclanthology.org/2005.mtsummit-papers.11
2005
Earlier work this paper cites.
Rosen A (2010) Morphological Tags in Parallel Corpora, pp 205–234
2010
Earlier work this paper cites.
Klementiev A, Titov I, Bhattarai B (2012) Inducing crosslingual distributed representations of words. In: Proceedings of COLING 2012, pp 1459–1474
2012
Earlier work this paper cites.
Tiedemann J (2012) Parallel data, tools and interfaces in OPUS. In: Chair) NCC, Choukri K, Declerck T, et al (eds) Proceedings of the Eight International Conference on Language Resources and Evaluation (LREC’12). European Language Resources Association (ELRA), Istanbul, Turkey
2012
Earlier work this paper cites.
De Marneffe MC, Dozat T, Silveira N, et al (2014) Universal stanford dependencies: A cross-linguistic typology. In: Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC’14), pp 4585–4592
2014
Earlier work this paper cites.
Mayer T, Cysouw M (2014) Creating a massively parallel Bible corpus. In: Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC’14). European Language Resources Association (ELRA), Reykjavik, Iceland, pp 3158–3163, URL http://www.lrec-conf.org/proceedings/lrec2014/pdf/220_Paper.pdf
2014
Cited alongside, same era.
Hammarström H (2015) Glottolog: A free, online, comprehensive bibliography of the world’s languages. In: 3rd International Conference on Linguistic and Cultural Diversity in Cyberspace, UNESCO, pp 183–188
2015
Cited alongside, same era.
Mogadala A, Rettinger A (2016) Bilingual word embeddings from parallel and non-parallel corpora for cross-language text classification. In: Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp 692–702
2016
Cited alongside, same era.
Conneau A, Khandelwal K, Goyal N, et al (2020) Unsupervised cross-lingual representation learning at scale. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Online, pp 8440–8451, 10.18653/v1/2020.acl-main.747 , URL https://aclanthology.org/2020.acl-main.747
2020
Later among the works it cites.
Joshi P, Santy S, Budhiraja A, et al (2020) The state and fate of linguistic diversity and inclusion in the NLP world. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Online, pp 6282–6293, 10.18653/v1/2020.acl-main.560 , URL https://aclanthology.org/2020.acl-main.560
2020
Later among the works it cites.
Lewis P, Oguz B, Rinott R, et al (2020) MLQA: Evaluating cross-lingual extractive question answering. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Online, pp 7315–7330, 10.18653/v1/2020.acl-main.653 , URL https://aclanthology.org/2020.acl-main.653
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rajpurkar P, Zhang J, Lopyrev K, et al (2016) SQuAD: 100,000+ questions for machine comprehension of text. In: Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Austin, Texas, pp 2383–2392, 10.18653/v1/D16-1264 , URL https://aclanthology.org/D16-1264
2016
Cited alongside, same era.
Ziemski M, Junczys-Dowmunt M, Pouliquen B (2016) The United Nations parallel corpus v1.0. In: Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16). European Language Resources Association (ELRA), Portorož, Slovenia, pp 3530–3534, URL https://aclanthology.org/L16-1561
2016
Cited alongside, same era.
2018
Cited alongside, same era.
Kunchukuttan A, Mehta P, Bhattacharyya P (2018) The IIT Bombay English-Hindi parallel corpus. In: Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018). European Language Resources Association (ELRA), Miyazaki, Japan, URL https://aclanthology.org/L18-1548
2018
Cited alongside, same era.
Schwenk H, Li X (2018) A corpus for multilingual document classification in eight languages. In: Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018). European Language Resources Association (ELRA), Miyazaki, Japan, URL https://aclanthology.org/L18-1560
2018
Cited alongside, same era.
Williams A, Nangia N, Bowman S (2018) A broad-coverage challenge corpus for sentence understanding through inference. In: Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). Association for Computational Linguistics, pp 1112–1122, URL http://aclweb.org/anthology/N18-1101
2018
Cited alongside, same era.
Agić Ž, Vulić I (2019) JW300: A wide-coverage parallel corpus for low-resource languages. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Florence, Italy, pp 3204–3210, 10.18653/v1/P19-1310 , URL https://aclanthology.org/P19-1310
2019
Cited alongside, same era.
Artetxe M, Schwenk H (2019) Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond. Transactions of the Association for Computational Linguistics 7:597–610. 10.1162/tacl_a_00288 , URL https://aclanthology.org/Q19-1038
2019
Cited alongside, same era.
Artetxe M, Ruder S, Yogatama D (2019) On the cross-lingual transferability of monolingual representations. arXiv preprint arXiv:191011856
2019
Cited alongside, same era.
Price I, Gifford-Moore J, Flemming J, et al (2020) Six attributes of unhealthy conversations. In: Proceedings of the Fourth Workshop on Online Abuse and Harms. Association for Computational Linguistics, Online, pp 114–124, 10.18653/v1/2020.alw-1.15 , URL https://aclanthology.org/2020.alw-1.15
2020
Later among the works it cites.
De Marneffe MC, Manning CD, Nivre J, et al (2021) Universal dependencies. Computational linguistics 47(2):255–308
2021
Later among the works it cites.
Ogueji K, Zhu Y, Lin J (2021) Small data? no problem! exploring the viability of pretrained multilingual language models for low-resourced languages. In: Proceedings of the 1st Workshop on Multilingual Representation Learning. Association for Computational Linguistics, Punta Cana, Dominican Republic, pp 116–126, 10.18653/v1/2021.mrl-1.11 , URL https://aclanthology.org/2021.mrl-1.11
2021
Later among the works it cites.
Adebara I, Elmadany A, Abdul-Mageed M, et al (2022) Serengeti: Massively multilingual language models for africa. 2212.10785
2022
Later among the works it cites.
Alabi JO, Adelani DI, Mosbach M, et al (2022) Adapting pre-trained language models to African languages via multilingual adaptive fine-tuning. In: Proceedings of the 29th International Conference on Computational Linguistics. International Committee on Computational Linguistics, Gyeongju, Republic of Korea, pp 4336–4349, URL https://aclanthology.org/2022.coling-1.382
2022
Later among the works it cites.
Nzeyimana A, Rubungo AN (2022) KinyaBERT: a morphology-aware kinyarwanda language model. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, 10.18653/v1/2022.acl-long.367 , URL https://doi.org/10.18653%2Fv1%2F2022.acl-long.367
2022
Later among the works it cites.
ImaniGooghari A, Lin P, Kargaran AH, et al (2023) Glot500: Scaling multilingual corpora and language models to 500 languages. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics
2023
Closest in time.
Petrov S, Das D, McDonald R (2012) A universal part-of-speech tagset. In: Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC’12). European Language Resources Association (ELRA), Istanbul, Turkey, pp 2089–2096, URL http://www.lrec-conf.org/proceedings/lrec2012/pdf/274_Paper.pdf
2096
Closest in time.