Fetching the paper…
Reading the bibliography…
Recent studies of large-scale contrastive pretraining in the text embedding domain show that using single-source minibatches, rather than mixed-source minibatches, can substantially improve overall model accuracy.
The Use of Hierarchic Clustering in Information Retrieval
N Jardine and C.J. van Rijsbergen · 1971
Earlier work this paper cites.
The cluster hypothesis revisited
Ellen M. Voorhees · 1985
Earlier work this paper cites.
Approximate nearest neighbor negative contrastive learning for dense text retrieval, 2020
Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul Bennett, Junaid Ahmed, and Arnold Overwijk · 2007
Earlier work this paper cites.
MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, Mir Rosenberg, Xia Song, Alina Stoica, Saurabh Tiwary, and Tong Wang · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Representation learning with contrastive predictive coding, 2019
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2019
Cited alongside, same era.
Efficiently teaching an effective dense retriever with balanced topic aware sampling, 2021
Sebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin, and Allan Hanbury · 2021
Cited alongside, same era.
Mteb: Massive text embedding benchmark, 2023
Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers · 2023
Cited alongside, same era.
Arctic-embed: Scalable, efficient, and accurate text embedding models, 2024
Luke Merrick, Danmei Xu, Gaurav Nuti, and Daniel Campos · 2024
Cited alongside, same era.
Nomic embed: Training a reproducible long context text embedder, 2024
Zach Nussbaum, John X. Morris, Brandon Duderstadt, and Andriy Mulyar · 2024
Closest in time.
The faiss library
Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou · 2024
Closest in time.
Text embeddings by weakly-supervised contrastive pre-training, 2024
Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei · 2024
Closest in time.
Mini-batch optimization of contrastive loss
Jaewoong Cho, Kartik Sreenivasan, Keon Lee, Kyunghoo Mun, Soheun Yi, Jeong-Gwan Lee, Anna Lee, Jy yong Sohn, Dimitris Papailiopoulos, and Kangwook Lee · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…