Fetching the paper…
Reading the bibliography…
While Indic NLP has made rapid advances recently in terms of the availability of corpora and pre-trained models, benchmark datasets on standard NLU tasks are limited.
Bleu: A method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalization
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020 · 2003
Earlier work this paper cites.
Spearman Rank Correlation Coefficient , pages 502–505. Springer New York, New York, NY
2008 · 2008
Earlier work this paper cites.
Pearson’s Correlation Coefficient , pages 1090–1091. Springer Netherlands, Dordrecht
Wilhelm Kirch, editor. 2008 · 2008
Earlier work this paper cites.
The back-translation score: Automatic mt evaluation at the sentence level without reference translations
Reinhard Rapp. 2009 · 2009
Earlier work this paper cites.
K-dpps: Fixed-size determinantal point processes
Alex Kulesza and Ben Taskar. 2011 · 2011
Earlier work this paper cites.
Determinantal point processes for machine learning
Alex Kulesza. 2012 · 2012
Earlier work this paper cites.
Recognizing textual entailment: Models and applications
Ido Dagan, Dan Roth, Mark Sammons, and Fabio Massimo Zanzotto. 2013 · 2013
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
An empirical investigation of catastrophic forgetting in gradient-based neural networks
Ian J. Goodfellow, Mehdi Mirza, Da Xiao, Aaron Courville, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Evaluation of the validity of back-translation as a method of assessing the accuracy of machine translation
Mai Miyabe and Takashi Yoshino. 2015 · 2015
Earlier work this paper cites.
SemEval-2016 task 1: Semantic textual similarity, monolingual and cross-lingual evaluation
Eneko Agirre, Carmen Banea, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, Rada Mihalcea, German Rigau, and Janyce Wiebe. 2016 · 2016
Earlier work this paper cites.
Assessing the use of back translation: the shortcomings of back translation as a quality testing method
Dorothée Behr. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
XNLI: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Earlier work this paper cites.
Second language acquisition and mother tongue influence of english language learners – a psycho analytic approach
A. Delbio, R. Abilasha, and m. Ilankumaran. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford and Karthik Narasimhan. 2018 · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
Cross-lingual language model pretraining
Alexis Conneau and Guillaume Lample. 2019 · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Cited alongside, same era.
Asynchronous pipelines for processing huge corpora on medium to low resource infrastructures
Pedro Javier Ortiz Su’arez, Benoit Sagot, and Laurent Romary. 2019 · 2019
Cited alongside, same era.
XGLUE: A new benchmark datasetfor cross-lingual pre-training, understanding and generation
Yaobo Liang, Nan Duan, Yeyun Gong, Ning Wu, Fenfei Guo, Weizhen Qi, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Ruofei Zhang, Rahul Agrawal, Edward Cui, Sining Wei, Taroon Bharti, Ying Qiao, Jiun-Hung Chen, Winnie Wu, Shuguang Liu, Fan Yang, Daniel Campos, Rangan Majumder, and Ming Zhou. 2020 · 2020
Later among the works it cites.
Multilingual denoising pre-training for neural machine translation
Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020 · 2020
Later among the works it cites.
A monolingual approach to contextualized word embeddings for mid-resource languages
Pedro Javier Ortiz Suárez, Laurent Romary, and Benoît Sagot. 2020 · 2020
Later among the works it cites.
Two-step classification using recasted data for low resource settings
Shagun Uppal, Vivek Gupta, Avinash Swaminathan, Haimin Zhang, Debanjan Mahata, Rakesh Gosangi, Rajiv Ratn Shah, and Amanda Stent. 2020 · 2020
Later among the works it cites.
CCNet: Extracting high quality monolingual datasets from web crawl data
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
How multilingual is multilingual BERT?
Telmo Pires, Eva Schlinger, and Dan Garrette. 2019 · 2019
Cited alongside, same era.
PAWS-X: A cross-lingual adversarial dataset for paraphrase identification
Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge. 2019 · 2019
Cited alongside, same era.
iNLTK: Natural language toolkit for indic languages
Gaurav Arora. 2020 · 2020
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Cited alongside, same era.
On the evaluation of machine translation systems trained with back-translation
Sergey Edunov, Myle Ott, Marc’Aurelio Ranzato, and Michael Auli. 2020 · 2020
Cited alongside, same era.
Statistical power and translationese in machine translation evaluation
Yvette Graham, Barry Haddow, and Philipp Koehn. 2020 · 2020
Cited alongside, same era.
Pmindia – a collection of parallel corpora of languages of india
Barry Haddow and Faheem Kirefu. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Later among the works it cites.
Muppet: Massive multi-task representations with pre-finetuning
Armen Aghajanyan, Anchit Gupta, Akshat Shrivastava, Xilun Chen, Luke Zettlemoyer, and Sonal Gupta. 2021 · 2021
Later among the works it cites.
Language invariant properties in natural language processing
Federico Bianchi, Debora Nozza, and Dirk Hovy. 2021 · 2021
Later among the works it cites.
Indicbart: A pre-trained model for natural language generation of indic languages
Raj Dabre, Himani Shrotriya, Anoop Kunchukuttan, Ratish Puduppully, Mitesh M. Khapra, and Pratyush Kumar. 2021 · 2021
Later among the works it cites.
Liro: Benchmark and leaderboard for romanian language tasks
Stefan Daniel Dumitrescu, Petru Rebeja, Beata Lorincz, Mihaela Gaman, Andrei Avram, Mihai Ilie, Andrei Pruteanu, Adriana Stan, Lorena Rosia, Cristina Iacobescu, Luciana Morogan, George Dima, Gabriel Marchidan, Traian Rebedea, Madalina Chitez, Dani Yogatama, Sebastian Ruder, Radu Tudor Ionescu, Razvan Pascanu, and Viorica Patraucean. 2021 · 2021
Later among the works it cites.
Analyzing the effects of reasoning types on cross-lingual transfer performance
Karthikeyan K, Aalok Sathe, Somak Aditya, and Monojit Choudhury. 2021 · 2021
Later among the works it cites.
Muril: Multilingual representations for indian languages
Simran Khanuja, Diksha Bansal, Sarvesh Mehtani, Savya Khosla, Atreyee Dey, Balaji Gopalan, Dilip Kumar Margam, Pooja Aggarwal, Rajiv Teja Nagipogu, Shachi Dave, Shruti Gupta, Subhash Chandra Bose Gali, Vish Subramanian, and Partha Talukdar. 2021 · 2021
Later among the works it cites.
Meta learning and its applications to natural language processing
Hung-yi Lee, Ngoc Thang Vu, and Shang-Wen Li. 2021 · 2021
Later among the works it cites.
Samanantar: The largest publicly available parallel corpora collection for 11 indic languages
Gowtham Ramesh, Sumanth Doddapaneni, Aravinth Bheemaraj, Mayank Jobanputra, Raghavan AK, Ajitesh Sharma, Sujit Sahoo, Harshita Diddee, Mahalakshmi J, Divyanshu Kakwani, Navneet Kumar, Aswin Pradeep, Kumar Deepak, Vivek Raghavan, Anoop Kunchukuttan, Pratyush Kumar, and Mitesh Shantadevi Khapra. 2021 · 2021
Later among the works it cites.
XTREME-R: Towards more challenging and nuanced multilingual evaluation
Sebastian Ruder, Noah Constant, Jan Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu, Junjie Hu, Dan Garrette, Graham Neubig, and Melvin Johnson. 2021 · 2021
Later among the works it cites.
mT5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021 · 2021
Later among the works it cites.