Fetching the paper…
Reading the bibliography…
In this paper, we introduce Neural Information Retrieval resources for 11 widely spoken Indian Languages (Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Oriya, Punjabi, Tamil, and Telugu) from two major Indian language families (Indo-Aryan and Dravidian).
The probabilistic relevance framework: Bm25 and beyond
Robertson, S., Zaragoza, H., et al · 2009
Earlier work this paper cites.
Overview of fire 2011
Palchowdhury, S., Majumder, P., Pal, D., Bandyopadhyay, A., and Mitra, M · 2013
Earlier work this paper cites.
Ms marco: A human generated machine reading comprehension dataset
Nguyen, T., Rosenberg, M., Song, X., Gao, J., Tiwary, S., Majumder, R., and Deng, L · 2016
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Natural questions: A benchmark for question answering research
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., Toutanova, K., Jones, L., Kelcey, M., Chang, M.-W., Dai, A. M., Uszkoreit, J., Le, Q., and Petrov, S · 2019
Earlier work this paper cites.
Cross-lingual language model pretraining
Lample, G., and Conneau, A · 2019
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering
Karpukhin, V., Oğuz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., and Yih, W.-t · 2020
Earlier work this paper cites.
The opennmt neural machine translation toolkit: 2020 edition
Klein, G., Hernandez, F., Nguyen, V., and Senellart, J · 2020
Earlier work this paper cites.
Clirmatrix: A massively large collection of bilingual and multilingual datasets for cross-lingual information retrieval
Sun, S., and Duh, K · 2020
Cited alongside, same era.
mmarco: A multilingual version of the ms marco passage ranking dataset
Bonifacio, L., Jeronymo, V., Abonizio, H. Q., Campiotti, I., Fadaee, M., Lotufo, R., and Nogueira, R · 2021
Cited alongside, same era.
Indicbart: A pre-trained model for indic natural language generation
Dabre, R., Shrotriya, H., Kunchukuttan, A., Puduppully, R., Khapra, M. M., and Kumar, P · 2021
Cited alongside, same era.
Colbertv2: Effective and efficient retrieval via lightweight late interaction
Santhanam, K., Khattab, O., Saad-Falcon, J., Potts, C., and Zaharia, M · 2021
Cited alongside, same era.
Beir: A heterogenous benchmark for zero-shot evaluation of information retrieval models
Samanantar: The largest publicly available parallel corpora collection for 11 indic languages
Ramesh, G., Doddapaneni, S., Bheemaraj, A., Jobanputra, M., Ak, R., Sharma, A., Sahoo, S., Diddee, H., Kakwani, D., Kumar, N., et al · 2022
Later among the works it cites.
Towards best practices for training multilingual dense retrieval models, 2022
Zhang, X., Ogueji, K., Ma, X., and Lin, J · 2022
Later among the works it cites.
Making a miracl: Multilingual information retrieval across a continuum of languages
Zhang, X., Thakur, N., Ogundepo, O., Kamalloo, E., Alfonso-Hermelo, D., Li, X., Liu, Q., Rezagholizadeh, M., and Lin, J · 2022
Later among the works it cites.
Cross-lingual knowledge transfer via distillation for multilingual information retrieval, 2023
Huang, Z., Yu, P., and Allan, J · 2023
Closest in time.
Improving cross-lingual information retrieval on low-resource languages via optimal transport distillation
Huang, Z., Yu, P., and Allan, J · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Thakur, N., Reimers, N., Rücklé, A., Srivastava, A., and Gurevych, I · 2021
Cited alongside, same era.
Mr. tydi: A multi-lingual benchmark for dense retrieval
Zhang, X., Ma, X., Shi, P., and Lin, J · 2021
Cited alongside, same era.
No language left behind: Scaling human-centered machine translation
Costa-jussà, M. R., Cross, J., Çelebi, O., Elbayad, M., Heafield, K., Heffernan, K., Kalbassi, E., Lam, J., Licht, D., Maillard, J., et al · 2022
Cited alongside, same era.
Closest in time.
Neural approaches to multilingual information retrieval
Lawrie, D., Yang, E., Oard, D. W., and Mayfield, J · 2023
Closest in time.
Simple yet effective neural ranking and reranking baselines for cross-lingual information retrieval
Lin, J., Alfonso-Hermelo, D., Jeronymo, V., Kamalloo, E., Lassance, C., Nogueira, R., Ogundepo, O., Rezagholizadeh, M., Thakur, N., Yang, J.-H., et al · 2023
Closest in time.