Fetching the paper…
Reading the bibliography…
Recently, embedding resources, including models, benchmarks, and datasets, have been widely released to support a variety of languages.
The merits of universal language model fine-tuning for small datasets–a case with dutch book reviews
Benjamin Van der Burgh and Suzan Verberne. 2019 · 1910
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1911
Earlier work this paper cites.
Wietse de Vries, Andreas van Cranenburgh, Arianna Bisazza, Tommaso Caselli, Gertjan van Noord, and Malvina Nissim. 2019 · 1912
Earlier work this paper cites.
The construction of a 500-million-word reference corpus of contemporary written dutch
Nelleke Oostdijk, Martin Reynaert, Véronique Hoste, and Ineke Schuurman. 2012 · 2012
Earlier work this paper cites.
Large scale syntactic annotation of written dutch: Lassy
Gertjan van Noord, Gosse Bouma, Frank Van Eynde, Daniël de Kok, Jelmer van der Linde, Ineke Schuurman, Erik Tjong Kim Sang, and Vincent Vandeghinste. 2012 · 2012
Earlier work this paper cites.
A sick cure for the evaluation of compositional distributional semantic models
Marco Marelli, Stefano Menini, Marco Baroni, Luisa Bentivogli, Raffaella Bernardi, and Roberto Zamparelli. 2014 · 2014
Earlier work this paper cites.
Ms marco: A human generated machine reading comprehension dataset
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, et al. 2016 · 2016
Earlier work this paper cites.
A full-text learning to rank dataset for medical information retrieval
Vera Boteva, Demian Gholipour, Artem Sokolov, and Stefan Riezler. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford and Karthik Narasimhan. 2018 · 2018
Earlier work this paper cites.
Fever: a large-scale dataset for fact extraction and verification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018 · 2018
Earlier work this paper cites.
Retrieval of the best counterargument without prior topic knowledge
Henning Wachsmuth, Shahbaz Syed, and Benno Stein. 2018 · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Natural questions: A benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019 · 2019
Earlier work this paper cites.
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Earlier work this paper cites.
SLURP: A spoken language understanding resource package
Emanuele Bastianelli, Andrea Vanzo, Pawel Swietojanski, and Verena Rieser. 2020 · 2020
Earlier work this paper cites.
SPECTER: Document-level representation learning using citation-informed transformers
Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel Weld. 2020 · 2020
Earlier work this paper cites.
RobBERT: a Dutch RoBERTa-based Language Model
Pieter Delobelle, Thomas Winters, and Bettina Berendt. 2020 · 2020
Earlier work this paper cites.
Xl-wic: A multilingual benchmark for evaluating semantic contextualization
A Raganato, T Pasini, J Camacho-Collados, M Pilehvar, et al. 2020 · 2020
Earlier work this paper cites.
Fact or fiction: Verifying scientific claims
David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020 · 2020
Earlier work this paper cites.
Fighting the COVID-19 infodemic: Modeling the perspective of journalists, fact-checkers, social media platforms, policy makers, and the society
Firoj Alam, Shaden Shaar, Fahim Dalvi, Hassan Sajjad, Alex Nikolov, Hamdy Mubarak, Giovanni Da San Martino, Ahmed Abdelali, Nadir Durrani, Kareem Darwish, Abdulaziz Al-Homaid, Wajdi Zaghouani, Tommaso Caselli, Gijs Danoe, Friso Stolk, Britt Bruntink, and Preslav Nakov. 2021 · 2021
Earlier work this paper cites.
mmarco: A multilingual version of ms marco passage ranking dataset
Luiz Henrique Bonifacio, Vitor Jeronymo, Hugo Queiroz Abonizio, Israel Campiotti, Marzieh Fadaee, , Roberto Lotufo, and Rodrigo Nogueira. 2021 · 2021
Earlier work this paper cites.
MultiEURLEX - a multi-lingual and multi-label legal document classification dataset for zero-shot cross-lingual transfer
Ilias Chalkidis, Manos Fergadiotis, and Ion Androutsopoulos. 2021 · 2021
Earlier work this paper cites.
SimCSE: Simple contrastive learning of sentence embeddings
Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021 · 2021
Earlier work this paper cites.
Machine translated multilingual sts benchmark dataset
Philip May. 2021 · 2021
Earlier work this paper cites.
BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021 · 2021
Earlier work this paper cites.
SICK-NL: A dataset for Dutch natural language inference
Gijs Wijnholds and Michael Moortgat. 2021 · 2021
Earlier work this paper cites.
Inpars: Unsupervised dataset generation for information retrieval
Luiz Bonifacio, Hugo Abonizio, Marzieh Fadaee, and Rodrigo Nogueira. 2022 · 2022
Earlier work this paper cites.
Domain- and task-adaptation for VaccinChatNL, a Dutch COVID-19 FAQ answering corpus and classification model
Jeska Buhmann, Maxime De Bruyn, Ehsan Lotfi, and Walter Daelemans. 2022 · 2022
Earlier work this paper cites.
No language left behind: Scaling human-centered machine translation
Marta R Costa-Jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, et al. 2022 · 2022
Cited alongside, same era.
Robbert-2022: Updating a dutch language model to account for evolving language use
Pieter Delobelle, Thomas Winters, and Bettina Berendt. 2022 · 2022
Cited alongside, same era.
Language-agnostic BERT sentence embedding
Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang. 2022 · 2022
Cited alongside, same era.
Unsupervised dense information retrieval with contrastive learning
Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2022 · 2022
Cited alongside, same era.
A statutory article retrieval dataset in French
Antoine Louis and Gerasimos Spanakis. 2022 · 2022
Cited alongside, same era.
Enhancing low-resource LLMs classification with PEFT and synthetic data
Parth Patwa, Simone Filice, Zhiyu Chen, Giuseppe Castellucci, Oleg Rokhlenko, and Shervin Malmasi. 2024 · 2024
Later among the works it cites.
Pl-mteb: Polish massive text embedding benchmark
Rafał Poświata, Sławomir Dadas, and Michał Perełkiewicz. 2024 · 2024
Later among the works it cites.
Attributed question answering for preconditions in the Dutch law
Felicia Redelaar, Romy Van Drie, Suzan Verberne, and Maaike De Boer. 2024 · 2024
Later among the works it cites.
Trans-tokenization and cross-lingual vocabulary transfers: Language adaptation of LLMs for low-resource NLP
François Remy, Pieter Delobelle, Hayastan Avetisyan, Alfiya Khabibullina, Miryam de Lhoneux, and Thomas Demeester. 2024 · 2024
Later among the works it cites.
Dan Saattrup Smart, Kenneth Enevoldsen, and Peter Schneider-Kamp. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Multilingual HateCheck: Functional tests for multilingual hate speech detection models
Paul Röttger, Haitham Seelawi, Debora Nozza, Zeerak Talat, and Bertie Vidgen. 2022 · 2022
Cited alongside, same era.
“zo grof !”: A comprehensive corpus for offensive and abusive language in Dutch
Ward Ruitenbeek, Victor Zwart, Robin Van Der Noord, Zhenja Gnezdilov, and Tommaso Caselli. 2022 · 2022
Cited alongside, same era.
Dutch news articles
Max Scheijen. 2022 · 2022
Cited alongside, same era.
Text embeddings by weakly-supervised contrastive pre-training
Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2022 · 2022
Cited alongside, same era.
Transfer learning for the visual arts: The multi-modal retrieval of iconclass codes
Nikolay Banar, Walter Daelemans, and Mike Kestemont. 2023 · 2023
Cited alongside, same era.
Dumb: A benchmark for smart evaluation of dutch models
Wietse de Vries, Martijn Wieling, and Malvina Nissim. 2023 · 2023
Cited alongside, same era.
Robbert-2023: Keeping dutch language models up-to-date at a lower cost thanks to model conversion
P Delobelle and F Remy. 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
The russian-focused embedders’ exploration: rumteb benchmark and russian embedding model design
Artem Snegirev, Maria Tikhonova, Anna Maksimova, Alena Fenogenova, and Alexander Abramov. 2024 · 2024
Later among the works it cites.
jina-embeddings-v3: Multilingual embeddings with task lora
Saba Sturua, Isabelle Mohr, Mohammad Kalim Akram, Michael Günther, Bo Wang, Markus Krimmel, Feng Wang, Georgios Mastrapas, Andreas Koukounas, Andreas Koukounas, Nan Wang, and Han Xiao. 2024 · 2024
Later among the works it cites.
Model2vec: Fast state-of-the-art static embeddings
Stephan Tulkens and Thomas van Dongen. 2024 · 2024
Later among the works it cites.
German text embedding clustering benchmark
Silvan Wehrli, Bert Arnrich, and Christopher Irrgang. 2024 · 2024
Later among the works it cites.
Beir-pl: Zero shot information retrieval benchmark for the polish language
Konrad Wojtasik, Kacper Wołowiec, Vadim Shishkin, Arkadiusz Janz, and Maciej Piasecki. 2024 · 2024
Later among the works it cites.
C-pack: Packed resources for general chinese embeddings
Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff, Defu Lian, and Jian-Yun Nie. 2024 · 2024
Later among the works it cites.
Arctic-embed 2.0: Multilingual retrieval without compromise
Puxuan Yu, Luke Merrick, Gaurav Nuti, and Daniel Campos. 2024 · 2024
Later among the works it cites.
mgte: Generalized long-context text representation and reranking models for multilingual text retrieval
Xin Zhang, Yanzhao Zhang, Dingkun Long, Wen Xie, Ziqi Dai, Jialong Tang, Huan Lin, Baosong Yang, Pengjun Xie, Fei Huang, et al. 2024 · 2024
Later among the works it cites.
Dense text retrieval based on pretrained language models: A survey
Wayne Xin Zhao, Jing Liu, Ruiyang Ren, and Ji-Rong Wen. 2024 · 2024
Later among the works it cites.
VABB-SHW: Dataset of Flemish Academic Bibliography for the Social Sciences and Humanities
Aspeslagh, P. and Guns, R. and Engels, T. C. E. 2024 · 2024
Later among the works it cites.
Dutch-CoLA (Revision 5a4196c)
Bylinina, Lisa and Abdi, Silvana and Brouwer, Hylke and Elzinga, Martine and Gunput, Shenza and Huisman, Sem and Krooneman, Collin and Poot, David and Top, Jelmer and Weideman, Cain. 2024 · 2024
Later among the works it cites.
Parul Awasthy, Aashka Trivedi, Yulong Li, Mihaela Bornea, David Cox, Abraham Daniels, Martin Franz, Gabe Goodhart, Bhavani Iyer, Vishwajeet Kumar, Luis Lastras, Scott McCarley, Rudra Murthy, Vignesh P, Sara Rosenthal, Salim Roukos, Jaydeep Sen, Sukriti Sharma, Avirup Sil, Kate Soule, Arafat Sultan, and Radu Florian. 2025 · 2025
Closest in time.
Swan and ArabicMTEB: Dialect-aware, Arabic-centric, cross-lingual, and cross-cultural embedding models and benchmarks
Gagan Bhatia, El Moatez Billah Nagoudi, Abdellah El Mekki, Fakhraddin Alwajih, and Muhammad Abdul-Mageed. 2025 · 2025
Closest in time.
Little giants: Synthesizing high-quality embedding data at scale
Haonan Chen, Liang Wang, Nan Yang, Yutao Zhu, Ziliang Zhao, Furu Wei, and Zhicheng Dou. 2025 · 2025
Closest in time.
Nv-retriever: Improving text embedding models with effective hard-negative mining
Gabriel de Souza P. Moreira, Radek Osmulski, Mengyao Xu, Ronay Ak, Benedikt Schifferer, and Even Oldridge. 2025 · 2025
Closest in time.
Detecting linguistic bias in government documents using large language models
Milena de Swart, Floris Den Hengst, and Jieying Chen. 2025 · 2025
Closest in time.
Webfaq: A multilingual collection of natural q&a datasets for dense retrieval
Michael Dinzinger, Laura Caspari, Kanishka Ghosh Dastidar, Jelena Mitrović, and Michael Granitzer. 2025 · 2025
Closest in time.
Kalm-embedding: Superior training data brings a stronger embedding model
Xinshuo Hu, Zifei Shan, Xinping Zhao, Zetian Sun, Zhenyu Liu, Dongfang Li, Shaolin Ye, Xinyuan Wei, Qian Chen, Baotian Hu, et al. 2025 · 2025
Closest in time.
Syntriever: How to train your retriever with synthetic data from LLMs
Minsang Kim and Seung Jun Baek. 2025 · 2025
Closest in time.
Nv-embed: Improved techniques for training llms as generalist embedding models
Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. 2025 · 2025
Closest in time.
Bilingual BSARD: Extending statutory article retrieval to Dutch
Ehsan Lotfi, Nikolay Banar, Nerses Yuzbashyan, and Walter Daelemans. 2025b · 2025
Closest in time.
Dutch news headlines for sarcasm detection
Harro Tuin. 2020 · 2025
Closest in time.
Qwen3 embedding: Advancing text embedding and reranking through foundation models
Yanzhao Zhang, Mingxin Li, Dingkun Long, Xin Zhang, Huan Lin, Baosong Yang, Pengjun Xie, An Yang, Dayiheng Liu, Junyang Lin, Fei Huang, and Jingren Zhou. 2025 · 2025
Closest in time.
Famteb: Massive text embedding benchmark in persian language
Erfan Zinvandi, Morteza Alikhani, Mehran Sarmadi, Zahra Pourbahman, Sepehr Arvin, Reza Kazemi, and Arash Amini. 2025 · 2025
Closest in time.
Mteb: Massive text embedding benchmark
Niklas Muennighoff, Nouamane Tazi, Loic Magne, and Nils Reimers. 2023 · 2037
Closest in time.