Fetching the paper…
Reading the bibliography…
In this report, we introduce Gemini Embedding, a state-of-the-art embedding model leveraging the power of Gemini, Google's most capable large language model.
Reciprocal rank fusion outperforms condorcet and individual rank learning methods
G. V. Cormack, C. L. Clarke, and S. Buettcher · 2009
Earlier work this paper cites.
Distributed representations of sentences and documents
Q. Le and T. Mikolov · 2014
Earlier work this paper cites.
Universal sentence encoder for english
D. Cer, Y. Yang, S.-y. Kong, N. Hua, N. Limtiaco, R. S. John, N. Constant, M. Guajardo-Cespedes, S. Yuan, C. Tar, et al · 2018
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
P. Izmailov, D. Podoprikhin, T. Garipov, D. Vetrov, and A. G. Wilson · 2018
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
J. Devlin, M. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Document expansion by query prediction
R. Nogueira, W. Yang, J. Lin, and K. Cho · 2019
Earlier work this paper cites.
Stochastic negative mining for learning with large output spaces
S. J. Reddi, S. Kale, F. Yu, D. Holtmann-Rice, J. Chen, and S. Kumar · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
N. Reimers and I. Gurevych · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering
V. Karpukhin, B. Oğuz, S. Min, P. Lewis, L. Y. Wu, S. Edunov, D. Chen, and W. tau Yih · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Earlier work this paper cites.
Xor qa: Cross-lingual open-retrieval question answering
A. Asai, J. Kasai, J. H. Clark, K. Lee, E. Choi, and H. Hajishirzi · 2021
Earlier work this paper cites.
Simcse: Simple contrastive learning of sentence embeddings
T. Gao, X. Yao, and D. Chen · 2021
Earlier work this paper cites.
Large dual encoders are generalizable retrievers
J. Ni, C. Qu, J. Lu, Z. Dai, G. H. ’Abrego, J. Ma, V. Zhao, Y. Luan, K. B. Hall, M.-W. Chang, and Y. Yang · 2021
Cited alongside, same era.
Inpars: Unsupervised dataset generation for information retrieval
L. Bonifacio, H. Abonizio, M. Fadaee, and R. Nogueira · 2022
Cited alongside, same era.
Promptagator: Few-shot dense retrieval from 8 examples
Z. Dai, V. Y. Zhao, J. Ma, Y. Luan, J. Ni, J. Lu, A. Bakalov, K. Guu, K. B. Hall, and M.-W. Chang · 2022
Cited alongside, same era.
Language-agnostic BERT sentence embedding
F. Feng, Y. Yang, D. Cer, N. Arivazhagan, and W. Wang · 2022
Cited alongside, same era.
Unsupervised dense information retrieval with contrastive learning
G. Izacard, M. Caron, L. Hosseini, S. Riedel, P. Bojanowski, A. Joulin, and E. Grave · 2022
Cited alongside, same era.
Matryoshka representation learning
Mteb: Massive text embedding benchmark
N. Muennighoff, N. Tazi, L. Magne, and N. Reimers · 2023
Later among the works it cites.
Xtreme-up: A user-centric scarce-data benchmark for under-represented languages
S. Ruder, J. H. Clark, A. Gutkin, M. Kale, M. Ma, M. Nicosia, S. Rijhwani, P. Riley, J.-M. Sarr, X. Wang, et al · 2023
Later among the works it cites.
Improving text embeddings with large language models
L. Wang, N. Yang, X. Huang, L. Yang, R. Majumder, and F. Wei · 2023
Later among the works it cites.
Miracl: A multilingual retrieval dataset covering 18 diverse languages
X. Zhang, N. Thakur, O. Ogundepo, E. Kamalloo, D. Alfonso-Hermelo, X. Li, Q. Liu, M. Rezagholizadeh, and J. Lin · 2023
Later among the works it cites.
Vlm2vec: Training vision-language models for massive multimodal embedding tasks
Z. Jiang, R. Meng, X. Yang, S. Yavuz, Y. Zhou, and W. Chen · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Kusupati, G. Bhatt, A. Rege, M. Wallingford, A. Sinha, V. Ramanujan, W. Howard-Snyder, K. Chen, S. Kakade, P. Jain, et al · 2022
Cited alongside, same era.
Text and code embeddings by contrastive pre-training
A. Neelakantan, T. Xu, R. Puri, A. Radford, J. M. Han, J. Tworek, Q. Yuan, N. Tezak, J. W. Kim, C. Hallacy, et al · 2022
Cited alongside, same era.
Sentence-t5: Scalable sentence encoders from pre-trained text-to-text models
J. Ni, G. H. Abrego, N. Constant, J. Ma, K. Hall, D. Cer, and Y. Yang · 2022
Cited alongside, same era.
Text embeddings by weakly-supervised contrastive pre-training
L. Wang, N. Yang, X. Huang, B. Jiao, L. Yang, D. Jiang, R. Majumder, and F. Wei · 2022
Cited alongside, same era.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
M. Wortsman, G. Ilharco, S. Y. Gadre, R. Roelofs, R. Gontijo-Lopes, A. S. Morcos, H. Namkoong, A. Farhadi, Y. Carmon, S. Kornblith, et al · 2022
Cited alongside, same era.
Inpars-v2: Large language models as efficient dataset generators for information retrieval
V. Jeronymo, L. Bonifacio, H. Abonizio, M. Fadaee, R. Lotufo, J. Zavrel, and R. Nogueira · 2023
Cited alongside, same era.
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M.-A. Lachaux, P. Stock, T. L. Scao, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed · 2023
Cited alongside, same era.
Gecko: Versatile text embeddings distilled from large language models
J. Lee, Z. Dai, X. Ren, B. Chen, D. Cer, J. R. Cole, K. Hui, M. Boratko, R. Kapadia, W. Ding, Y. Luan, S. M. K. Duddu, G. H. Abrego, W. Shi, N. Gupta, A. Kusupati, P. Jain, S. R. Jonnalagadda, M.-W. Chang, and I. Naim · 2024
Later among the works it cites.
Making text embedders few-shot learners
C. Li, M. Qin, S. Xiao, J. Chen, K. Luo, Y. Shao, D. Lian, and Z. Liu · 2024
Later among the works it cites.
Sfrembedding-mistral: enhance text retrieval with transfer learning
R. Meng, Y. Liu, S. R. Joty, C. Xiong, Y. Zhou, and S. Yavuz · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024
G. Team · 2024
Later among the works it cites.
Leveraging LLMs for synthesizing training data across many languages in multilingual dense retrieval
N. Thakur, J. Ni, G. Hernandez Abrego, J. Wieting, J. Lin, and D. Cer · 2024
Later among the works it cites.
Mmteb: Massive multilingual text embedding benchmark
K. Enevoldsen, I. Chung, I. Kerboua, M. Kardos, A. Mathur, D. Stap, J. Gala, W. Siblini, D. Krzemiński, G. I. Winata, et al · 2025
Closest in time.
Nv-embed: Improved techniques for training llms as generalist embedding models
C. Lee, R. Roy, M. Xu, J. Raiman, M. Shoeybi, B. Catanzaro, and W. Ping · 2025
Closest in time.
Adapting decoder-based language models for diverse encoder downstream tasks, 2025
P. Suganthan, F. Moiseev, L. Yan, J. Wu, J. Ni, J. Han, I. Zitouni, E. Alfonseca, X. Wang, and Z. Dong · 2025
Closest in time.