Fetching the paper…
Reading the bibliography…
In this paper, we introduce a comprehensive benchmark for Persian (Farsi) text embeddings, built upon the Massive Text Embedding Benchmark (MTEB).
Dense passage retrieval for open-domain question answering
Vladimir Karpukhin, Barlas Oğuz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020 · 2004
Earlier work this paper cites.
Deepsentipers: Novel deep learning models trained over proposed augmented persian sentiment corpus
Javad Pourmostafa Roshan Sharami, Parsa Abbasi Sarabestani, and Seyed Abolghasem Mirroshandel. 2020 · 2004
Earlier work this paper cites.
Farstail: A persian natural language inference dataset
Hossein Amirkhani, Mohammad AzariJafari, Zohreh Pourjafari, Soroush Faridan-Jahromi, Zeinab Kouhkan, and Azadeh Amirak. 2020 · 2009
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
A SICK cure for the evaluation of compositional distributional semantic models
Marco Marelli, Stefano Menini, Marco Baroni, Luisa Bentivogli, Raffaella Bernardi, and Roberto Zamparelli. 2014 · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
MS MARCO: A human generated machine reading comprehension dataset
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016 · 2016
Earlier work this paper cites.
SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Earlier work this paper cites.
Sentipers: A sentiment analysis corpus for persian
Pedram Hosseini, Ali Ahmadian Ramaki, Hassan Maleki, Mansoureh Anvari, and Seyed Abolghasem Mirroshandel. 2018 · 2018
Earlier work this paper cites.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych. 2019b · 2019
Earlier work this paper cites.
community-datasets/farsi_news
Mehdi Allahyar. 2020 · 2020
Earlier work this paper cites.
Parsbert: Transformer-based model for persian language understanding
Mehrdad Farahani, Mohammad Gharachorloo, Marzieh Farahani, and Mohammad Manthouri. 2020 · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020 · 2020
Cited alongside, same era.
mmarco: A multilingual version of MS MARCO passage ranking dataset
Luiz Henrique Bonifacio, Israel Campiotti, Roberto A. Lotufo, and Rodrigo Frassetto Nogueira. 2021 · 2021
Cited alongside, same era.
Farsick: A persian semantic textual similarity and natural language inference dataset
Zahra Ghasemi and Mohammad Ali Keyvanrad. 2021 · 2021
Cited alongside, same era.
ParsiNLU: A suite of language understanding challenges for Persian
Daniel Khashabi, Arman Cohan, Siamak Shakeri, Pedram Hosseini, Pouya Pezeshkpour, Malihe Alikhani, Moin Aminnaseri, Marzieh Bitaab, Faeze Brahman, Sarik Ghazarian, Mozhdeh Gheini, Arman Kabiri, Rabeeh Karimi Mahabagdi, Omid Memarrast, Ahmadreza Mosallanezhad, Erfan Noury, Shahab Raji, Mohammad Sadegh Rasooli, Sepideh Sadeghi, Erfan Sadeqi Azer, Niloofar Safi Samghabadi, Mahsa Shafaei, Saber Sheybani, Ali Tazarv, and Yadollah Yaghoobzadeh. 2021 · 2021
Cited alongside, same era.
Towards general text embeddings with multi-stage contrastive learning
Zehan Li, Xin Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, and Meishan Zhang. 2023 · 2023
Later among the works it cites.
Hezar: The all-in-one ai library for persian
Aryan Shekarlaban and Pooya Mohammadi Kazaj. 2023 · 2023
Later among the works it cites.
MIRACL: A multilingual retrieval dataset covering 18 diverse languages
Xinyu Zhang, Nandan Thakur, Odunayo Ogundepo, Ehsan Kamalloo, David Alfonso-Hermelo, Xiaoguang Li, Qun Liu, Mehdi Rezagholizadeh, and Jimmy Lin. 2023 · 2023
Later among the works it cites.
M3-embedding: Multi-linguality, multi-functionality, multi-granularity text embeddings through self-knowledge distillation
Jianlyu Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, and Zheng Liu. 2024 · 2024
Later among the works it cites.
Mteb-french: Resources for french sentence embedding evaluation and analysis
Mathieu Ciancone, Imene Kerboua, Marion Schaeffer, and Wissam Siblini. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nandan Thakur, Nils Reimers, Andreas Ruckl’e, Abhishek Srivastava, and Iryna Gurevych. 2021 · 2021
Cited alongside, same era.
Language-agnostic BERT sentence embedding
Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang. 2022 · 2022
Cited alongside, same era.
alighasemi/farsi_paraphrase_detection
Ali Ghasemi. 2022 · 2022
Cited alongside, same era.
Seyedali/persian-text-emotion
Seyed Ali Mir Mohammad Hosseini. 2022 · 2022
Cited alongside, same era.
Overview of the trec 2022 neuclir track
Dawn J Lawrie, Sean MacAvaney, James Mayfield, Paul McNamee, Douglas W. Oard, Luca Soldaini, and Eugene Yang. 2023 · 2022
Cited alongside, same era.
Exappc: a large-scale persian paraphrase detection corpus
Reyhaneh Sadeghi, Hamed Karbasi, and Ahmad Akbari. 2022 · 2022
Cited alongside, same era.
Text embeddings by weakly-supervised contrastive pre-training
Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2022 · 2022
Cited alongside, same era.
hamedhf/nlp_twitter_analysis
Hamed Feizabadi. 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
Fabert: Pre-training bert on persian blogs
Mostafa Masumi, Seyed Soroush Majd, Mehrnoush Shamsfard, and Hamid Beigy. 2024 · 2024
Later among the works it cites.
Pl-mteb: Polish massive text embedding benchmark
Rafał Poświata, Sławomir Dadas, and Michał Perełkiewicz. 2024 · 2024
Later among the works it cites.
Tookabert: A step forward for persian nlu
MohammadAli SadraeiJavaheri, Ali Moghaddaszadeh, Milad Molazadeh, Fariba Naeiji, Farnaz Aghababaloo, Hamideh Rafiee, Zahra Amirmahani, Tohid Abedini, Fatemeh Zahra Sheikhi, and Amirmohammad Salehoof. 2024 · 2024
Later among the works it cites.
jina-embeddings-v3: Multilingual embeddings with task lora
Saba Sturua, Isabelle Mohr, Mohammad Kalim Akram, Michael Günther, Bo Wang, Markus Krimmel, Feng Wang, Georgios Mastrapas, Andreas Koukounas, Andreas Koukounas, Nan Wang, and Han Xiao. 2024 · 2024
Later among the works it cites.
Multilingual e5 text embeddings: A technical report
Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. 2024 · 2024
Later among the works it cites.
C-pack: Packed resources for general chinese embeddings
Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff, Defu Lian, and Jian-Yun Nie. 2024 · 2024
Later among the works it cites.
mGTE: Generalized long-context text representation and reranking models for multilingual text retrieval
Xin Zhang, Yanzhao Zhang, Dingkun Long, Wen Xie, Ziqi Dai, Jialong Tang, Huan Lin, Baosong Yang, Pengjun Xie, Fei Huang, Meishan Zhang, Wenjie Li, and Min Zhang. 2024 · 2024
Later among the works it cites.
Persian web document retrieval corpus
Erfan Zinvandi, Morteza Alikhani, Zahra Pourbahman, Reza Kazemi, and Arash Amini. 2024 · 2024
Later among the works it cites.
MTEB: Massive text embedding benchmark
Niklas Muennighoff, Nouamane Tazi, Loic Magne, and Nils Reimers. 2023 · 2037
Closest in time.