Fetching the paper…
Reading the bibliography…
In this paper, we introduce the Polish Massive Text Embedding Benchmark (PL-MTEB), a comprehensive benchmark for text embeddings in the Polish language.
V-Measure: A Conditional Entropy-Based External Cluster Evaluation Measure
Andrew Rosenberg and Julia Hirschberg. 2007 · 2007
Earlier work this paper cites.
A Survey of Text Clustering Algorithms , pages 77–128
Charu C. Aggarwal and ChengXiang Zhai. 2012 · 2012
Earlier work this paper cites.
Efficient Estimation of Word Representations in Vector Space
Tomas Mikolov, Kai Chen, and Greg Corrado ands Jeffrey Dean. 2013a · 2013
Earlier work this paper cites.
A SICK cure for the evaluation of compositional distributional semantic models
Marco Marelli, Stefano Menini, Marco Baroni, Luisa Bentivogli, Raffaella Bernardi, and Roberto Zamparelli. 2014 · 2014
Earlier work this paper cites.
The Polish Summaries Corpus
Maciej Ogrodniczuk and Mateusz Kopeć. 2014 · 2014
Earlier work this paper cites.
GloVe: Global Vectors for Word Representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
Daniel Fernando Campos, Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, Li Deng, and Bhaskar Mitra. 2016 · 2016
Earlier work this paper cites.
Enriching Word Vectors with Subword Information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Polish evaluation dataset for compositional distributional semantics models
Alina Wróblewska and Katarzyna Krasnowska-Kieraś. 2017 · 2017
Earlier work this paper cites.
SentEval: An evaluation toolkit for universal sentence representations
Alexis Conneau and Douwe Kiela. 2018 · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Multi-level sentiment analysis of PolEmo 2.0: Extended corpus of multi-domain consumer reviews
Jan Kocoń, Piotr Miłkowski, and Monika Zaśko-Zielińska. 2019 · 2019
Earlier work this paper cites.
Empirical Linguistic Study of Sentence Embeddings
Katarzyna Krasnowska-Kieraś and Alina Wróblewska. 2019 · 2019
Earlier work this paper cites.
Results of the PolEval 2019 Shared Task 6: First Dataset and Open Shared Task for Automatic Cyberbullying Detection in Polish Twitter
Michal Ptaszynski, Agata Pieciukiewicz, and Paweł Dybała. 2019 · 2019
Earlier work this paper cites.
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Earlier work this paper cites.
SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Earlier work this paper cites.
Pre-training Polish Transformer-Based Language Models at Scale
Sławomir Dadas, Michał Perełkiewicz, and Rafał Poświata. 2020b · 2020
Earlier work this paper cites.
Embedding-based retrieval in facebook search
Jui-Ting Huang, Ashish Sharma, Shuying Sun, Li Xia, David Zhang, Philip Pronin, Janani Padmanabhan, Giuseppe Ottaviano, and Linjun Yang. 2020 · 2020
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020 · 2020
Cited alongside, same era.
Making monolingual sentence embeddings multilingual using knowledge distillation
Nils Reimers and Iryna Gurevych. 2020 · 2020
Cited alongside, same era.
KLEJ: Comprehensive Benchmark for Polish Language Understanding
Piotr Rybak, Robert Mroczkowski, Janusz Tracz, and Ireneusz Gawlik. 2020 · 2020
Cited alongside, same era.
SimCSE: Simple contrastive learning of sentence embeddings
Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021 · 2021
Cited alongside, same era.
Machine translated multilingual sts benchmark dataset
Philip May. 2021 · 2021
Cited alongside, same era.
HerBERT: Efficiently Pretrained Transformer-based Language Model for Polish
Robert Mroczkowski, Piotr Rybak, Alina Wróblewska, and Ireneusz Gawlik. 2021 · 2021
Gemma 2: Improving open language models at a practical size
Gemma Team, Google DeepMind. 2024 · 2024
Closest in time.
Silver Retriever: Advancing Neural Passage Retrieval for Polish Question Answering
Piotr Rybak and Maciej Ogrodniczuk. 2024 · 2024
Closest in time.
BEIR-PL: Zero shot information retrieval benchmark for the Polish language
Konrad Wojtasik, Kacper Wołowiec, Vadim Shishkin, Arkadiusz Janz, and Maciej Piasecki. 2024 · 2024
Closest in time.
C-pack: Packed resources for general chinese embeddings
Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff, Defu Lian, and Jian-Yun Nie. 2024 · 2024
Closest in time.
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, and 39 others. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021 · 2021
Cited alongside, same era.
TSDAE: Using transformer-based sequential denoising auto-encoderfor unsupervised sentence embedding learning
Kexin Wang, Nils Reimers, and Iryna Gurevych. 2021 · 2021
Cited alongside, same era.
This is the way: designing and compiling LEPISZCZE, a comprehensive NLP benchmark for Polish
Lukasz Augustyniak, Kamil Tagowski, Albert Sawczyn, Denis Janiak, Roman Bartusiak, Adrian Szymczak, Arkadiusz Janz, Piotr Szymański, Marcin Wątroba, Mikoł aj Morzy, Tomasz Kajdanowicz, and Maciej Piasecki. 2022 · 2022
Cited alongside, same era.
Training Effective Neural Sentence Encoders from Automatically Mined Paraphrases
Sławomir Dadas. 2022 · 2022
Cited alongside, same era.
Language-agnostic BERT sentence embedding
Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang. 2022 · 2022
Cited alongside, same era.
Jack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie, Kay Rottmann, Ana Sanchez, Aaron Nash, Liam Urbach, Vishesh Kakarala, Richa Singh, Swetha Ranganath, Laurie Crist, Misha Britan, Wouter Leeuwis, Gokhan Tur, and Prem Natarajan. 2022 · 2022
Cited alongside, same era.
Closest in time.
Arctic-embed 2.0: Multilingual retrieval without compromise
Puxuan Yu, Luke Merrick, Gaurav Nuti, and Daniel Campos. 2024 · 2024
Closest in time.
mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval
Xin Zhang, Yanzhao Zhang, Dingkun Long, Wen Xie, Ziqi Dai, Jialong Tang, Huan Lin, Baosong Yang, Pengjun Xie, Fei Huang, and 1 others. 2024 · 2024
Closest in time.
Mteb-nl and e5-nl: Embedding benchmark and models for dutch
Nikolay Banar, Ehsan Lotfi, Jens Van Nooten, Cristina Arhiliuc, Marija Kliocaite, and Walter Daelemans. 2025 · 2025
Closest in time.
TR-MTEB: A comprehensive benchmark and embedding model suite for Turkish sentence representations
Mehmet Selman Baysan and Tunga Gungor. 2025 · 2025
Closest in time.
Swan and ArabicMTEB: Dialect-aware, Arabic-centric, cross-lingual, and cross-cultural embedding models and benchmarks
Gagan Bhatia, El Moatez Billah Nagoudi, Abdellah El Mekki, Fakhraddin Alwajih, and Muhammad Abdul-Mageed. 2025 · 2025
Closest in time.
Mmteb: Massive multilingual text embedding benchmark
Kenneth Enevoldsen, Isaac Chung, Imene Kerboua, Márton Kardos, Ashwin Mathur, David Stap, Jay Gala, Wissam Siblini, Dominik Krzeminski, {Genta Indra} Winata, Saba Sturua, Saiteja Utpala, Mathieu Ciancone, Marion Schaeffer, Gabriel Sequeira, Diganta Misra, Shreeya Dhakal, Jonathan Rystrøm, Roman Solomatin, and 37 others. 2025 · 2025
Closest in time.
KaLM-Embedding: Superior Training Data Brings A Stronger Embedding Model
Xinshuo Hu, Zifei Shan, Xinping Zhao, Zetian Sun, Zhenyu Liu, Dongfang Li, Shaolin Ye, Xinyuan Wei, Qian Chen, Baotian Hu, and 1 others. 2025 · 2025
Closest in time.
DRAMA: Diverse augmentation from large language models to smaller dense retrievers
Xueguang Ma, Xi Victoria Lin, Barlas Oguz, Jimmy Lin, Wen-tau Yih, and Xilun Chen. 2025 · 2025
Closest in time.
Vn-mteb: Vietnamese massive text embedding benchmark
Loc Pham, Tung Luu, Thu Vo, Minh Nguyen, and Viet Hoang. 2025 · 2025
Closest in time.
The Russian-focused embedders’ exploration: ruMTEB benchmark and Russian embedding model design
Artem Snegirev, Maria Tikhonova, Maksimova Anna, Alena Fenogenova, and Aleksandr Abramov. 2025 · 2025
Closest in time.
Afrimteb and afrie5: Benchmarking and adapting text embedding models for african languages
Kosei Uemura, Miaoran Zhang, and David Ifeoluwa Adelani. 2025 · 2025
Closest in time.
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, and 41 others. 2025 · 2025
Closest in time.
FaMTEB: Massive text embedding benchmark in Persian language
Erfan Zinvandi, Morteza Alikhani, Mehran Sarmadi, Zahra Pourbahman, Sepehr Arvin, Reza Kazemi, and Arash Amini. 2025 · 2025
Closest in time.
MTEB: Massive Text Embedding Benchmark
Niklas Muennighoff, Nouamane Tazi, Loic Magne, and Nils Reimers. 2023 · 2037
Closest in time.