Fetching the paper…
Reading the bibliography…
We introduce RusBEIR, a comprehensive benchmark designed for zero-shot evaluation of information retrieval (IR) models in the Russian language.
“Russian Information Retrieval Evaluation Seminar.”
Boris Dobrov et al · 2004
Earlier work this paper cites.
“Morphological analyzer and generator for Russian and Ukrainian languages”
Mikhail Korobov · 2015
Earlier work this paper cites.
“chrF: character n-gram F-score for automatic MT evaluation”
Maja Popović · 2015
Earlier work this paper cites.
“Ms marco: A human generated machine reading comprehension dataset”
Payal Bajaj et al · 2016
Earlier work this paper cites.
“Natural language processing: python and NLTK”
Nitin Hardeniya et al · 2016
Earlier work this paper cites.
“Squad: 100,000+ questions for machine comprehension of text”
P Rajpurkar · 2016
Earlier work this paper cites.
“BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2019
Earlier work this paper cites.
“Self-attentive model for headline generation”
Daniil Gavrilov, Pavel Kalaidin and Valentin Malykh · 2019
Earlier work this paper cites.
“On the Cross-lingual Transferability of Monolingual Representations”
Mikel Artetxe, Sebastian Ruder and Dani Yogatama · 2020
Earlier work this paper cites.
“Tydi qa: A benchmark for information-seeking question answering in ty pologically di verse languages”
Jonathan Clark et al · 2020
Earlier work this paper cites.
“Sberquad–russian reading comprehension dataset: Description and analysis”
Pavel Efimov, Andrey Chertok, Leonid Boytsov and Pavel Braslavski · 2020
Cited alongside, same era.
“mmarco: A multilingual version of the ms marco passage ranking dataset”
Luiz Bonifacio et al · 2021
Cited alongside, same era.
“RuBQ 2.0: an innovated Russian question answering dataset”
Ivan Rybin, Vladislav Korablinov, Pavel Efimov and Pavel Braslavski · 2021
Cited alongside, same era.
“BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models”
Nandan Thakur et al · 2021
Cited alongside, same era.
“Language-agnostic BERT Sentence Embedding”
Fangxiaoyu Feng et al · 2022
Cited alongside, same era.
“Text embeddings by weakly-supervised contrastive pre-training”
“Hindi-BEIR: A Large Scale Retrieval Benchmark in Hindi”
Arkadeep Acharya, Rudra Murthy, Vishwajeet Kumar and Jaydeep Sen · 2024
Later among the works it cites.
“BEIR-NL: Zero-shot Information Retrieval Benchmark for the Dutch Language”
Nikolay Banar, Ehsan Lotfi and Walter Daelemans · 2024
Later among the works it cites.
Jianlv Chen et al · 2024
Later among the works it cites.
“PIRB: A Comprehensive Benchmark of Polish Dense and Hybrid Text Retrieval Methods”
Slawomir Dadas, Michał Perełkiewicz and Rafał Poświata · 2024
Later among the works it cites.
“The Russian-focused embedders’ exploration: ruMTEB benchmark and Russian embedding model design”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liang Wang et al · 2022
Cited alongside, same era.
“Proceedings of the 15th Annual Meeting of the Forum for Information Retrieval Evaluation”
Debasis Ganguly et al · 2023
Cited alongside, same era.
“PolEval 2022/23 challenge tasks and results”
Łukasz Kobyliński et al · 2023
Cited alongside, same era.
“Fact-checking benchmark for the Russian Large Language Models”
Anastasia Kozlova, Denis Shevelev and Alena Fenogenova · 2023
Cited alongside, same era.
“Miracl: A multilingual retrieval dataset covering 18 diverse languages”
Xinyu Zhang et al · 2023
Cited alongside, same era.
Artem Snegirev et al · 2024
Later among the works it cites.
“Multilingual e5 text embeddings: A technical report”
Liang Wang et al · 2024
Later among the works it cites.
“BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language”
Konrad Wojtasik et al · 2024
Later among the works it cites.
“A Family of Pretrained Transformer Language Models for Russian”
Dmitry Zmitrovich et al · 2024
Later among the works it cites.
“MTEB: Massive Text Embedding Benchmark”
Niklas Muennighoff, Nouamane Tazi, Loic Magne and Nils Reimers · 2037
Closest in time.