Fetching the paper…
Reading the bibliography…
We introduce EmbeddingGemma, a new lightweight, open text embedding model based on the Gemma 3 language model family.
Learning spread-out local feature descriptors
X. Zhang, F. X. Yu, S. Kumar, and S. Chang · 2017
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
P. Izmailov, D. Podoprikhin, T. Garipov, D. P. Vetrov, and A. G. Wilson · 2018
Earlier work this paper cites.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko · 2018
Earlier work this paper cites.
Xor qa: Cross-lingual open-retrieval question answering
A. Asai, J. Kasai, J. H. Clark, K. Lee, E. Choi, and H. Hajishirzi · 2021
Earlier work this paper cites.
Towards unsupervised dense information retrieval with contrastive learning
G. Izacard, M. Caron, L. Hosseini, S. Riedel, P. Bojanowski, A. Joulin, and E. Grave · 2021
Earlier work this paper cites.
Large dual encoders are generalizable retrievers, 2021
J. Ni, C. Qu, J. Lu, Z. Dai, G. H. Ábrego, J. Ma, V. Y. Zhao, Y. Luan, K. B. Hall, M.-W. Chang, and Y. Yang · 2021
Earlier work this paper cites.
Simcse: Simple contrastive learning of sentence embeddings, 2022
T. Gao, X. Yao, and D. Chen · 2022
Earlier work this paper cites.
Matryoshka representation learning
A. Kusupati, G. Bhatt, A. Rege, M. Wallingford, A. Sinha, V. Ramanujan, W. Howard-Snyder, K. Chen, S. Kakade, P. Jain, and A. Farhadi · 2022
Earlier work this paper cites.
Mteb: Massive text embedding benchmark
N. Muennighoff, N. Tazi, L. Magne, and N. Reimers · 2022
Earlier work this paper cites.
Text and code embeddings by contrastive pre-training, 2022
A. Neelakantan, T. Xu, R. Puri, A. Radford, J. M. Han, J. Tworek, Q. Yuan, N. Tezak, J. W. Kim, C. Hallacy, J. Heidecke, P. Shyam, B. Power, T. E. Nekoul, G. Sastry, G. Krueger, D. Schnurr, F. P. Such, K. Hsu, M. Thompson, T. Khan, T. Sherbakov, J. Jang, P. Welinder, and L. Weng · 2022
Earlier work this paper cites.
Colbertv2: Effective and efficient retrieval via lightweight late interaction, 2022
K. Santhanam, O. Khattab, J. Saad-Falcon, C. Potts, and M. Zaharia · 2022
Cited alongside, same era.
M. Wortsman, G. Ilharco, S. Y. Gadre, R. Roelofs, R. Gontijo-Lopes, A. S. Morcos, H. Namkoong, A. Farhadi, Y. Carmon, S. Kornblith, and L. Schmidt · 2022
Cited alongside, same era.
Embeddistill: A geometric knowledge distillation for information retrieval, 2023
S. Kim, A. S. Rawat, M. Zaheer, S. Jayasumana, V. Sadhanala, W. Jitkrittum, A. K. Menon, R. Fergus, and S. Kumar · 2023
Cited alongside, same era.
Xtreme-up: A user-centric scarce-data benchmark for under-represented languages
S. Ruder, J. H. Clark, A. Gutkin, M. Kale, M. Ma, M. Nicosia, S. Rijhwani, P. Riley, J.-M. Sarr, X. Wang, et al · 2023
Cited alongside, same era.
Longembed: Extending embedding models for long context retrieval, 2024
D. Zhu, L. Wang, N. Yang, Y. Song, W. Wu, F. Wei, and S. Li · 2024
Later among the works it cites.
Small language models are the future of agentic ai, 2025
P. Belcak, G. Heinrich, S. Diao, Y. Fu, X. Dong, S. Muralidharan, Y. C. Lin, and P. Molchanov · 2025
Closest in time.
Nv-retriever: Improving text embedding models with effective hard-negative mining, 2025
G. de Souza P. Moreira, R. Osmulski, M. Xu, R. Ak, B. Schifferer, and E. Oldridge · 2025
Closest in time.
Mmteb: Massive multilingual text embedding benchmark
K. Enevoldsen, I. Chung, I. Kerboua, M. Kardos, A. Mathur, D. Stap, J. Gala, W. Siblini, D. Krzemiński, G. I. Winata, et al · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Su, W. Shi, J. Kasai, Y. Wang, Y. Hu, M. Ostendorf, W. tau Yih, N. A. Smith, L. Zettlemoyer, and T. Yu · 2023
Cited alongside, same era.
Ul2: Unifying language learning paradigms, 2023
Y. Tay, M. Dehghani, V. Q. Tran, X. Garcia, J. Wei, X. Wang, H. W. Chung, S. Shakeri, D. Bahri, T. Schuster, H. S. Zheng, D. Zhou, N. Houlsby, and D. Metzler · 2023
Cited alongside, same era.
E5-v: Universal embeddings with multimodal large language models, 2024
T. Jiang, M. Song, Z. Zhang, H. Huang, W. Deng, F. Sun, Q. Zhang, D. Wang, and F. Zhuang · 2024
Cited alongside, same era.
Gecko: Versatile text embeddings distilled from large language models
J. Lee, Z. Dai, X. Ren, B. Chen, D. Cer, J. R. Cole, K. Hui, M. Boratko, R. Kapadia, W. Ding, Y. Luan, S. M. K. Duddu, G. H. Abrego, W. Shi, N. Gupta, A. Kusupati, P. Jain, S. R. Jonnalagadda, M.-W. Chang, and I. Naim · 2024
Cited alongside, same era.
jina-embeddings-v3: Multilingual embeddings with task lora, 2024
S. Sturua, I. Mohr, M. K. Akram, M. Günther, B. Wang, M. Krimmel, F. Wang, G. Mastrapas, A. Koukounas, N. Wang, and H. Xiao · 2024
Cited alongside, same era.
Nv-embed: Improved techniques for training llms as generalist embedding models, 2025a
C. Lee, R. Roy, M. Xu, J. Raiman, M. Shoeybi, B. Catanzaro, and W. Ping
Cited in the paper.
Gemini embedding: Generalizable embeddings from gemini, 2025b
J. Lee, F. Chen, S. Dua, D. Cer, M. Shanbhogue, I. Naim, G. H. Ábrego, Z. Li, K. Chen, H. S. Vera, X. Ren, S. Zhang, D. Salz, M. Boratko, J. Han, B. Chen, S. Huang, V. Rao, P. Suganthan, F. Han, A. Doumanoglou, N. Gupta, F. Moiseev, C. Yip, A. Jain, S. Baumgartner, S. Shahi, F. P. Gomez, S. Mariserla, M. Choi, P. Shah, S. Goenka, K. Chen, Y. Xia, K. Chen, S. M. K. Duddu, Y. Chen, T. Walker, W. Zhou, R. Ghiya, Z. Gleicher, K. Gill, Z. Dong, M. Seyedhosseini, Y. Sung, R. Hoffmann, and T. Duerig
Cited in the paper.
Text embeddings by weakly-supervised contrastive pre-training, 2024a
L. Wang, N. Yang, X. Huang, B. Jiao, L. Yang, D. Jiang, R. Majumder, and F. Wei
Cited in the paper.
Z. Jiang, R. Meng, X. Yang, S. Yavuz, Y. Zhou, and W. Chen · 2025
Closest in time.
Llave: Large language and vision embedding models with hardness-weighted contrastive learning, 2025
Z. Lan, L. Niu, F. Meng, J. Zhou, and J. Su · 2025
Closest in time.
Generative representational instruction tuning, 2025
N. Muennighoff, H. Su, L. Wang, N. Yang, F. Wei, T. Yu, A. Singh, and D. Kiela · 2025
Closest in time.
Adapting decoder-based language models for diverse encoder downstream tasks, 2025
P. Suganthan, F. Moiseev, L. Yan, J. Wu, J. Ni, J. Han, I. Zitouni, E. Alfonseca, X. Wang, and Z. Dong · 2025
Closest in time.
Gemma 3 technical report, 2025
G. Team · 2025
Closest in time.