Fetching the paper…
Reading the bibliography…
We present Gecko, a compact and versatile text embedding model.
Reciprocal rank fusion outperforms condorcet and individual rank learning methods
G. V. Cormack, C. L. Clarke, and S. Buettcher · 2009
Earlier work this paper cites.
Distributed representations of sentences and documents
Q. Le and T. Mikolov · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
S. Bowman, G. Angeli, C. Potts, and C. D. Manning · 2015
Earlier work this paper cites.
Universal sentence encoder for english
D. Cer, Y. Yang, S.-y. Kong, N. Hua, N. Limtiaco, R. S. John, N. Constant, M. Guajardo-Cespedes, S. Yuan, C. Tar, et al · 2018
Earlier work this paper cites.
Fever: a large-scale dataset for fact extraction and verification
J. Thorne, A. Vlachos, C. Christodoulopoulos, and A. Mittal · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
A. Williams, N. Nangia, and S. R. Bowman · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. Cohen, R. Salakhutdinov, and C. D. Manning · 2018
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, et al · 2019
Earlier work this paper cites.
Latent retrieval for weakly supervised open domain question answering
K. Lee, M.-W. Chang, and K. Toutanova · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
N. Reimers and I. Gurevych · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering
V. Karpukhin, B. Oğuz, S. Min, P. Lewis, L. Y. Wu, S. Edunov, D. Chen, and W. tau Yih · 2020
Earlier work this paper cites.
Approximate nearest neighbor negative contrastive learning for dense text retrieval
L. Xiong, C. Xiong, Y. Li, K.-F. Tang, J. Liu, P. Bennett, J. Ahmed, and A. Overwijk · 2020
Earlier work this paper cites.
Simcse: Simple contrastive learning of sentence embeddings
T. Gao, X. Yao, and D. Chen · 2021
Earlier work this paper cites.
Distilling knowledge from reader to retriever for question answering
G. Izacard and E. Grave · 2021
Earlier work this paper cites.
Learning dense representations of phrases at scale
J. Lee, M. Sung, J. Kang, and D. Chen · 2021
Earlier work this paper cites.
Large dual encoders are generalizable retrievers
J. Ni, C. Qu, J. Lu, Z. Dai, G. H. ’Abrego, J. Ma, V. Zhao, Y. Luan, K. B. Hall, M.-W. Chang, and Y. Yang · 2021
Cited alongside, same era.
Rocketqa: An optimized training approach to dense passage retrieval for open-domain question answering
Y. Qu, Y. Ding, J. Liu, K. Liu, R. Ren, W. X. Zhao, D. Dong, H. Wu, and H. Wang · 2021
Cited alongside, same era.
RocketQAv2: A joint training method for dense passage retrieval and passage re-ranking
R. Ren, Y. Qu, J. Liu, W. X. Zhao, Q. She, H. Wu, H. Wang, and J.-R. Wen · 2021
Cited alongside, same era.
Beir: A heterogeneous benchmark for zero-shot evaluation of information retrieval models
N. Thakur, N. Reimers, A. Rücklé, A. Srivastava, and I. Gurevych · 2021
Cited alongside, same era.
mt5: A massively multilingual pre-trained text-to-text transformer
L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel · 2021
Cited alongside, same era.
Text embeddings by weakly-supervised contrastive pre-training
L. Wang, N. Yang, X. Huang, B. Jiao, L. Yang, D. Jiang, R. Majumder, and F. Wei · 2022
Later among the works it cites.
R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen, et al · 2023
Later among the works it cites.
PaRaDe: Passage ranking using demonstrations with LLMs
A. Drozdov, H. Zhuang, Z. Dai, Z. Qin, R. Rahimi, X. Wang, D. Alon, M. Iyyer, A. McCallum, D. Metzler, and K. Hui · 2023
Later among the works it cites.
Inpars-v2: Large language models as efficient dataset generators for information retrieval
V. Jeronymo, L. Bonifacio, H. Abonizio, M. Fadaee, R. Lotufo, J. Zavrel, and R. Nogueira · 2023
Later among the works it cites.
Towards general text embeddings with multi-stage contrastive learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Asai, T. Schick, P. Lewis, X. Chen, G. Izacard, S. Riedel, H. Hajishirzi, and W.-t. Yih · 2022
Cited alongside, same era.
Inpars: Data augmentation for information retrieval using large language models
L. Bonifacio, H. Abonizio, M. Fadaee, and R. Nogueira · 2022
Cited alongside, same era.
Promptagator: Few-shot dense retrieval from 8 examples
Z. Dai, V. Y. Zhao, J. Ma, Y. Luan, J. Ni, J. Lu, A. Bakalov, K. Guu, K. B. Hall, and M.-W. Chang · 2022
Cited alongside, same era.
Unsupervised dense information retrieval with contrastive learning
G. Izacard, M. Caron, L. Hosseini, S. Riedel, P. Bojanowski, A. Joulin, and E. Grave · 2022
Cited alongside, same era.
Matryoshka representation learning
A. Kusupati, G. Bhatt, A. Rege, M. Wallingford, A. Sinha, V. Ramanujan, W. Howard-Snyder, K. Chen, S. Kakade, P. Jain, et al · 2022
Cited alongside, same era.
Text and code embeddings by contrastive pre-training
A. Neelakantan, T. Xu, R. Puri, A. Radford, J. M. Han, J. Tworek, Q. Yuan, N. Tezak, J. W. Kim, C. Hallacy, et al · 2022
Cited alongside, same era.
Sentence-t5: Scalable sentence encoders from pre-trained text-to-text models
J. Ni, G. H. Abrego, N. Constant, J. Ma, K. Hall, D. Cer, and Y. Yang · 2022
Cited alongside, same era.
Z. Li, X. Zhang, Y. Zhang, D. Long, P. Xie, and M. Zhang · 2023
Later among the works it cites.
Fine-tuning llama for multi-stage text retrieval
X. Ma, L. Wang, N. Yang, F. Wei, and J. Lin · 2023
Later among the works it cites.
Samtone: Improving contrastive loss for dual encoder retrieval models with same tower negatives
F. Moiseev, G. H. Abrego, P. Dornbach, I. Zitouni, E. Alfonseca, and Z. Dong · 2023
Later among the works it cites.
Mteb: Massive text embedding benchmark
N. Muennighoff, N. Tazi, L. Magne, and N. Reimers · 2023
Later among the works it cites.
Questions are all you need to train a dense passage retriever
D. S. Sachan, M. Lewis, D. Yogatama, L. Zettlemoyer, J. Pineau, and M. Zaheer · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
G. Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2023
Later among the works it cites.
Improving text embeddings with large language models
L. Wang, N. Yang, X. Huang, L. Yang, R. Majumder, and F. Wei · 2023
Later among the works it cites.
Miracl: A multilingual retrieval dataset covering 18 diverse languages
X. Zhang, N. Thakur, O. Ogundepo, E. Kamalloo, D. Alfonso-Hermelo, X. Li, Q. Liu, M. Rezagholizadeh, and J. Lin · 2023
Later among the works it cites.
Beyond yes and no: Improving zero-shot llm rankers via scoring fine-grained relevance labels
H. Zhuang, Z. Qin, K. Hui, J. Wu, L. Yan, X. Wang, and M. Berdersky · 2023
Later among the works it cites.
Leveraging llms for unsupervised dense retriever ranking
E. Khramtsova, S. Zhuang, M. Baktashmotlagh, and G. Zuccon · 2024
Closest in time.
Generative representational instruction tuning
N. Muennighoff, H. Su, L. Wang, N. Yang, F. Wei, T. Yu, A. Singh, and D. Kiela · 2024
Closest in time.
Repetition improves language model embeddings
J. M. Springer, S. Kotha, D. Fried, G. Neubig, and A. Raghunathan · 2024
Closest in time.