Fetching the paper…
Reading the bibliography…
Script diversity presents a challenge to Multilingual Language Models (MLLM) by reducing lexical overlap among closely related languages.
Similarity of neural network representations revisited
Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey E. Hinton. 2019 · 1905
Earlier work this paper cites.
Individual comparisons by ranking methods
Frank Wilcoxon. 1945 · 1945
Earlier work this paper cites.
On a Test of Whether one of Two Random Variables is Stochastically Larger than the Other
H. B. Mann and D. R. Whitney. 1947 · 1947
Earlier work this paper cites.
Passage to india? anuradhapura and the early use of the brahmi script
R.A.E. Coningham, F.R. Allchin, C.M. Batt, and D. Lucy. 1996 · 1996
Earlier work this paper cites.
The world's writing systems
Charles F. Hockett, Peter T. Daniels, and William Bright. 1997 · 1997
Earlier work this paper cites.
Rethinking embedding coupling in pre-trained language models
Hyung Won Chung, Thibault Févry, Henry Tsai, Melvin Johnson, and Sebastian Ruder. 2020 · 2010
Earlier work this paper cites.
Using Effect Size-or Why the P Value Is Not Enough
G. M. Sullivan and R. Feinn. 2012 · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Aspect based sentiment analysis: Category detection and sentiment classification for hindi
Md. Shad Akhtar, Asif Ekbal, and Pushpak Bhattacharyya. 2016 · 2016
Earlier work this paper cites.
Phonologically aware neural model for named entity recognition in low resource transfer settings
Akash Bharadwaj, David Mortensen, Chris Dyer, and Jaime Carbonell. 2016 · 2016
Earlier work this paper cites.
An empirical study of language relatedness for transfer learning in neural machine translation
Raj Dabre, Tetsuji Nakagawa, and Hideto Kazawa. 2017 · 2017
Earlier work this paper cites.
URIEL and lang2vec: Representing languages as typological, geographical, and phylogenetic vectors
Patrick Littell, David R. Mortensen, Ke Lin, Katherine Kairis, Carlisle Turner, and Lori Levin. 2017 · 2017
Earlier work this paper cites.
ACTSA: Annotated corpus for Telugu sentiment analysis
Sandeep Sricharan Mukku and Radhika Mamidi. 2017 · 2017
Earlier work this paper cites.
SVCCA: singular vector canonical correlation analysis for deep learning dynamics and interpretability
Maithra Raghu, Justin Gilmer, Jason Yosinski, and Jascha Sohl-Dickstein. 2017 · 2017
Earlier work this paper cites.
Adapting word embeddings to new languages with morphological and phonological subword representations
Aditi Chaudhary, Chunting Zhou, Lori Levin, Graham Neubig, David R. Mortensen, and Jaime Carbonell. 2018 · 2018
Earlier work this paper cites.
Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Earlier work this paper cites.
Insights on representational similarity in neural networks with canonical correlation
Ari S. Morcos, Maithra Raghu, and Samy Bengio. 2018 · 2018
Earlier work this paper cites.
Exploring BERT’s Vocabulary
Judit Ács. 2019 · 2019
Earlier work this paper cites.
Pushing the limits of low-resource morphological inflection
Antonios Anastasopoulos and Graham Neubig. 2019 · 2019
Earlier work this paper cites.
Investigating multilingual NMT representations at scale
Sneha Kudugunta, Ankur Bapna, Isaac Caswell, and Orhan Firat. 2019 · 2019
Earlier work this paper cites.
Asynchronous pipelines for processing huge corpora on medium to low resource infrastructures
Pedro Javier Ortiz Su’arez, Benoit Sagot, and Laurent Romary. 2019 · 2019
Cited alongside, same era.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Cited alongside, same era.
How multilingual is multilingual BERT?
Telmo Pires, Eva Schlinger, and Dan Garrette. 2019 · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019 · 2019
Cited alongside, same era.
Massively multilingual transfer for NER
Afshin Rahimi, Yuan Li, and Trevor Cohn. 2019 · 2019
Cited alongside, same era.
Why You Should Do NLP Beyond English
Sebastian Ruder. 2020 · 2020
Later among the works it cites.
Pre-training via leveraging assisting languages for neural machine translation
Haiyue Song, Raj Dabre, Zhuoyuan Mao, Fei Cheng, Sadao Kurohashi, and Eiichiro Sumita. 2020 · 2020
Later among the works it cites.
SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python
Pauli Virtanen, Ralf Gommers, Travis E. Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, Stéfan J. van der Walt, Matthew Brett, Joshua Wilson, K. Jarrod Millman, Nikolay Mayorov, Andrew R. J. Nelson, Eric Jones, Robert Kern, Eric Larson, C J Carey, İlhan Polat, Yu Feng, Eric W. Moore, Jake VanderPlas, Denis Laxalde, Josef Perktold, Robert Cimrman, Ian Henriksen, E. A. Quintero, Charles R. Harris, Anne M. Archibald, Antônio H. Ribeiro, Fabian Pedregosa, Paul van Mulbregt, and SciPy 1.0 Contributors. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shruti Rijhwani, Jiateng Xie, Graham Neubig, and Jaime G. Carbonell. 2019 · 2019
Cited alongside, same era.
BERT is not an interlingua and the bias of tokenization
Jasdeep Singh, Bryan McCann, Richard Socher, and Caiming Xiong. 2019 · 2019
Cited alongside, same era.
On Romanization for model transfer between scripts in neural machine translation
Chantal Amrhein and Rico Sennrich. 2020 · 2020
Cited alongside, same era.
On the cross-lingual transferability of monolingual representations
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020 · 2020
Cited alongside, same era.
Emerging cross-lingual structure in pretrained language models
Alexis Conneau, Shijie Wu, Haoran Li, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Cited alongside, same era.
An annotated dataset of discourse modes in Hindi stories
Swapnil Dhanwal, Hritwik Dutta, Hitesh Nankani, Nilay Shrivastava, Yaman Kumar, Junyi Jessy Li, Debanjan Mahata, Rakesh Gosangi, Haimin Zhang, Rajiv Ratn Shah, and Amanda Stent. 2020 · 2020
Cited alongside, same era.
Efficient neural machine translation for low-resource languages via exploiting related languages
Vikrant Goyal, Sourav Kumar, and Dipti Misra Sharma. 2020 · 2020
Cited alongside, same era.
Shijie Wu and Mark Dredze. 2020 · 2020
Later among the works it cites.
Ungoliant: An optimized pipeline for the generation of a very large-scale multilingual web corpus
Julien Abadji, Pedro Javier Ortiz Suárez, Laurent Romary, and Benoît Sagot. 2021 · 2021
Later among the works it cites.
Ecco: An open source library for the explainability of transformer language models
J Alammar. 2021 · 2021
Later among the works it cites.
Establishing interlingua in multilingual language models
Maksym Del and Mark Fishel. 2021 · 2021
Later among the works it cites.
Role of Language Relatedness in Multilingual Fine-tuning of Language Models: A Case Study in Indo-Aryan Languages
Tejas Dhamecha, Rudra Murthy, Samarth Bharadwaj, Karthik Sankaranarayanan, and Pushpak Bhattacharyya. 2021 · 2021
Later among the works it cites.
Exploiting language relatedness for low web-resource language model adaptation: An Indic languages study
Yash Khemchandani, Sarvesh Mehtani, Vaidehi Patil, Abhijeet Awasthi, Partha Talukdar, and Sunita Sarawagi. 2021 · 2021
Later among the works it cites.
Datasets: A community library for natural language processing
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, and Thomas Wolf. 2021 · 2021
Later among the works it cites.
On the stability of fine-tuning BERT: misconceptions, explanations, and strong baselines
Marius Mosbach, Maksym Andriushchenko, and Dietrich Klakow. 2021 · 2021
Later among the works it cites.
When being unseen from mBERT is just the beginning: Handling new languages with multilingual language models
Benjamin Muller, Antonios Anastasopoulos, Benoît Sagot, and Djamé Seddah. 2021 · 2021
Later among the works it cites.
First align, then predict: Understanding the cross-lingual ability of multilingual BERT
Benjamin Müller, Yanai Elazar, Benoît Sagot, and Djamé Seddah. 2021 · 2021
Later among the works it cites.
Unks everywhere: Adapting multilingual language models to new scripts
Jonas Pfeiffer, Ivan Vulic, Iryna Gurevych, and Sebastian Ruder. 2021 · 2021
Later among the works it cites.
Aksharamukha transliteration tool
Vinodh Rajan. 2015 · 2021
Later among the works it cites.
How good is your tokenizer? on the monolingual performance of multilingual language models
Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, and Iryna Gurevych. 2021 · 2021
Later among the works it cites.
The Flores-101 evaluation benchmark for low-resource and multilingual machine translation
Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzmán, and Angela Fan. 2022 · 2022
Closest in time.
Pyicu transliteration tool
PyICU · 2022
Closest in time.