Fetching the paper…
Reading the bibliography…
In this paper, we systematically study the potential of pre-training with Large Language Model(LLM)-based document expansion for dense passage retrieval.
Document Expansion by Query Prediction
Nogueira, R. F.; Yang, W.; Lin, J.; and Cho, K. 2019 · 1904
Earlier work this paper cites.
Overview of the TREC 2019 deep learning track
Craswell, N.; Mitra, B.; Yilmaz, E.; Campos, D.; and Voorhees, E. M. 2020 · 2003
Earlier work this paper cites.
Curriculum learning
Bengio, Y.; Louradour, J.; Collobert, R.; and Weston, J. 2009 · 2009
Earlier work this paper cites.
The probabilistic relevance framework: BM25 and beyond
Robertson, S.; Zaragoza, H.; et al. 2009 · 2009
Earlier work this paper cites.
MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
Nguyen, T.; Rosenberg, M.; Song, X.; Gao, J.; Tiwary, S.; Majumder, R.; and Deng, L. 2016 · 2016
Earlier work this paper cites.
Relevance-Based Language Models
Lavrenko, V.; and Croft, W. B. 2017 · 2017
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Earlier work this paper cites.
FAQ Retrieval using Query-Question Similarity and BERT-Based Query-Answer Relevance
Sakata, W.; Shibata, T.; Tanaka, R.; and Kurohashi, S. 2019 · 2019
Earlier work this paper cites.
Pre-training Tasks for Embedding-based Large-scale Retrieval
Chang, W.; Yu, F. X.; Chang, Y.; Yang, Y.; and Kumar, S. 2020 · 2020
Earlier work this paper cites.
Overview of the TREC 2020 deep learning track
Craswell, N.; Mitra, B.; Yilmaz, E.; and Campos, D. 2021 · 2020
Earlier work this paper cites.
Dense Passage Retrieval for Open-Domain Question Answering
Karpukhin, V.; Oguz, B.; Min, S.; Lewis, P.; Wu, L.; Edunov, S.; Chen, D.; and Yih, W.-t. 2020 · 2020
Earlier work this paper cites.
ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
Khattab, O.; and Zaharia, M. 2020 · 2020
Earlier work this paper cites.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Lewis, P. S. H.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Küttler, H.; Lewis, M.; Yih, W.; Rocktäschel, T.; Riedel, S.; and Kiela, D. 2020 · 2020
Earlier work this paper cites.
CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data
Wenzek, G.; Lachaux, M.; Conneau, A.; Chaudhary, V.; Guzmán, F.; Joulin, A.; and Grave, E. 2020 · 2020
Earlier work this paper cites.
Condenser: a Pre-training Architecture for Dense Retrieval
Gao, L.; and Callan, J. 2021 · 2021
Earlier work this paper cites.
SimCSE: Simple Contrastive Learning of Sentence Embeddings
Gao, T.; Yao, X.; and Chen, D. 2021 · 2021
Cited alongside, same era.
Towards Unsupervised Dense Information Retrieval with Contrastive Learning
Izacard, G.; Caron, M.; Hosseini, L.; Riedel, S.; Bojanowski, P.; Joulin, A.; and Grave, E. 2021 · 2021
Cited alongside, same era.
Pre-trained Language Model for Web-scale Retrieval in Baidu Search
Liu, Y.; Lu, W.; Cheng, S.; Shi, D.; Wang, S.; Cheng, Z.; and Yin, D. 2021 · 2021
Cited alongside, same era.
Less is More: Pretrain a Strong Siamese Encoder for Dense Text Retrieval Using a Weak Decoder
Lu, S.; He, D.; Xiong, C.; Ke, G.; Malik, W.; Dou, Z.; Bennett, P.; Liu, T.-Y.; and Overwijk, A. 2021 · 2021
Cited alongside, same era.
RocketQA: An Optimized Training Approach to Dense Passage Retrieval for Open-Domain Question Answering
Qu, Y.; Ding, Y.; Liu, J.; Liu, K.; Ren, R.; Zhao, W. X.; Dong, D.; Wu, H.; and Wang, H. 2021 · 2021
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C. L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P. F.; Leike, J.; and Lowe, R. 2022 · 2022
Later among the works it cites.
ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction
Santhanam, K.; Khattab, O.; Saad-Falcon, J.; Potts, C.; and Zaharia, M. 2022 · 2022
Later among the works it cites.
Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks
Wang, Y.; Mishra, S.; Alipoormolabashi, P.; Kordi, Y.; Mirzaei, A.; Naik, A.; Ashok, A.; Dhanasekaran, A. S.; Arunkumar, A.; Stap, D.; Pathak, E.; Karamanolakis, G.; Lai, H. G.; Purohit, I.; Mondal, I.; Anderson, J.; Kuznia, K.; Doshi, K.; Pal, K. K.; Patel, M.; Moradshahi, M.; Parmar, M.; Purohit, M.; Varshney, N.; Kaza, P. R.; Verma, P.; Puri, R. S.; Karia, R.; Doshi, S.; Sampat, S. K.; Mishra, S.; A, S. R.; Patro, S.; Dixit, T.; and Shen, X. 2022b · 2022
Later among the works it cites.
Query-as-context Pre-training for Dense Passage Retrieval
Wu, X.; Ma, G.; and Hu, S. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
RocketQAv2: A Joint Training Method for Dense Passage Retrieval and Passage Re-ranking
Ren, R.; Qu, Y.; Liu, J.; Zhao, W. X.; She, Q.; Wu, H.; Wang, H.; and Wen, J.-R. 2021 · 2021
Cited alongside, same era.
BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models
Thakur, N.; Reimers, N.; Rücklé, A.; Srivastava, A.; and Gurevych, I. 2021 · 2021
Cited alongside, same era.
Recent advances in retrieval-augmented text generation
Cai, D.; Wang, Y.; Liu, L.; and Shi, S. 2022 · 2022
Cited alongside, same era.
Query Generation with External Knowledge for Dense Retrieval
Cho, S.; Jeong, S.; Yang, W.; and Park, J. C. 2022 · 2022
Cited alongside, same era.
PaLM: Scaling Language Modeling with Pathways
Chowdhery, A.; Narang, S.; Devlin, J.; Bosma, M.; Mishra, G.; Roberts, A.; Barham, P.; Chung, H. W.; Sutton, C.; Gehrmann, S.; Schuh, P.; Shi, K.; Tsvyashchenko, S.; Maynez, J.; Rao, A.; Barnes, P.; Tay, Y.; Shazeer, N.; Prabhakaran, V.; Reif, E.; Du, N.; Hutchinson, B.; Pope, R.; Bradbury, J.; Austin, J.; Isard, M.; Gur-Ari, G.; Yin, P.; Duke, T.; Levskaya, A.; Ghemawat, S.; Dev, S.; Michalewski, H.; Garcia, X.; Misra, V.; Robinson, K.; Fedus, L.; Zhou, D.; Ippolito, D.; Luan, D.; Lim, H.; Zoph, B.; Spiridonov, A.; Sepassi, R.; Dohan, D.; Agrawal, S.; Omernick, M.; Dai, A. M.; Pillai, T. S.; Pellat, M.; Lewkowycz, A.; Moreira, E.; Child, R.; Polozov, O.; Lee, K.; Zhou, Z.; Wang, X.; Saeta, B.; Diaz, M.; Firat, O.; Catasta, M.; Wei, J.; Meier-Hellstern, K.; Eck, D.; Dean, J.; Petrov, S.; and Fiedel, N. 2022 · 2022
Cited alongside, same era.
Unsupervised Corpus Aware Language Model Pre-training for Dense Passage Retrieval
Gao, L.; and Callan, J. 2022 · 2022
Cited alongside, same era.
Tevatron: An efficient and flexible toolkit for dense retrieval
Gao, L.; Ma, X.; Lin, J.; and Callan, J. 2022 · 2022
Cited alongside, same era.
Zhou, K.; Liu, X.; Gong, Y.; Zhao, W. X.; Jiang, D.; Duan, N.; and Wen, J.-R. 2022 · 2022
Later among the works it cites.
Precise Zero-Shot Dense Retrieval without Relevance Labels
Gao, L.; Ma, X.; Lin, J.; and Callan, J. 2023 · 2023
Closest in time.
Query Expansion by Prompting Large Language Models
Jagerman, R.; Zhuang, H.; Qin, Z.; Wang, X.; and Bendersky, M. 2023 · 2023
Closest in time.
RetroMAE-2: Duplex Masked Auto-Encoder For Pre-Training Retrieval-Oriented Language Models
Liu, Z.; Xiao, S.; Shao, Y.; and Cao, Z. 2023 · 2023
Closest in time.
LLaMA: Open and Efficient Foundation Language Models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; Rodriguez, A.; Joulin, A.; Grave, E.; and Lample, G. 2023 · 2023
Closest in time.
Query2doc: Query Expansion with Large Language Models
Wang, L.; Yang, N.; and Wei, F. 2023 · 2023
Closest in time.
Self-Instruct: Aligning Language Models with Self-Generated Instructions
Wang, Y.; Kordi, Y.; Mishra, S.; Liu, A.; Smith, N. A.; Khashabi, D.; and Hajishirzi, H. 2023 · 2023
Closest in time.
ConTextual Masked Auto-Encoder for Dense Passage Retrieval
Wu, X.; Ma, G.; Lin, M.; Lin, Z.; Wang, Z.; and Hu, S. 2023a · 2023
Closest in time.
Generate rather than Retrieve: Large Language Models are Strong Context Generators
Yu, W.; Iter, D.; Wang, S.; Xu, Y.; Ju, M.; Sanyal, S.; Zhu, C.; Zeng, M.; and Jiang, M. 2023 · 2023
Closest in time.
Pre-trained Language Model-based Retrieval and Ranking for Web Search
Zou, L.; Lu, W.; Liu, Y.; Cai, H.; Chu, X.; Ma, D.; Shi, D.; Sun, Y.; Cheng, Z.; Gu, S.; Wang, S.; and Yin, D. 2023 · 2023
Closest in time.