Fetching the paper…
Reading the bibliography…
Large language models (LLMs) typically utilize the top-k contexts from a retriever in retrieval-augmented generation (RAG).
Simple bm25 extension to multiple weighted fields
Robertson, S., Zaragoza, H., and Taylor, M · 2004
Earlier work this paper cites.
Semantic parsing on freebase from question-answer pairs
Berant, J., Chou, A., Frostig, R., and Liang, P · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
An overview of the bioasq large-scale biomedical semantic indexing and question answering competition
Tsatsaronis, G., Balikas, G., Malakasiotis, P., Partalas, I., Zschunke, M., Alvers, M. R., Weissenborn, D., Krithara, A., Petridis, S., Polychronopoulos, D., et al · 2015
Earlier work this paper cites.
Ms marco: A human generated machine reading comprehension dataset
Bajaj, P., Campos, D., Craswell, N., Deng, L., Gao, J., Liu, X., Majumder, R., McNamara, A., Mitra, B., Nguyen, T., et al · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Joshi, M., Choi, E., Weld, D., and Zettlemoyer, L · 2017
Earlier work this paper cites.
Newsqa: A machine comprehension dataset
Trischler, A., Wang, T., Yuan, X., Harris, J., Sordoni, A., Bachman, P., and Suleman, K · 2017
Earlier work this paper cites.
The narrativeqa reading comprehension challenge
Kočiskỳ, T., Schwarz, J., Blunsom, P., Dyer, C., Hermann, K. M., Melis, G., and Grefenstette, E · 2018
Earlier work this paper cites.
An introduction to neural information retrieval
Mitra, B., Craswell, N., et al · 2018
Earlier work this paper cites.
Fever: A large-scale dataset for fact extraction and verification
Thorne, J., Vlachos, A., Christodoulopoulos, C., and Mittal, A · 2018
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W. W., Salakhutdinov, R., and Manning, C. D · 2018
Earlier work this paper cites.
Quoref: A reading comprehension dataset with questions requiring coreferential reasoning
Dasigi, P., Liu, N. F., Marasović, A., Smith, N. A., and Gardner, M · 2019
Earlier work this paper cites.
Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dua, D., Wang, Y., Dasigi, P., Stanovsky, G., Singh, S., and Gardner, M · 2019
Earlier work this paper cites.
Eli5: Long form question answering
Fan, A., Jernite, Y., Perez, E., Grangier, D., Weston, J., and Auli, M · 2019
Earlier work this paper cites.
Pubmedqa: A dataset for biomedical research question answering
Jin, Q., Dhingra, B., Liu, Z., Cohen, W., and Lu, X · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., et al · 2019
Earlier work this paper cites.
Reasoning over paragraph effects in situations
Lin, K., Tafjord, O., Clark, P., and Gardner, M · 2019
Earlier work this paper cites.
doc2dial: A goal-oriented document-grounded dialogue dataset
Feng, S., Wan, H., Gunasekara, C., Patel, S., Joshi, S., and Lastras, L · 2020
Earlier work this paper cites.
Retrieval augmented language model pre-training
Guu, K., Lee, K., Tung, Z., Pasupat, P., and Chang, M · 2020
Earlier work this paper cites.
Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps
Ho, X., Nguyen, A.-K. D., Sugawara, S., and Aizawa, A · 2020
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering
Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., and Yih, W.-t · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al · 2020
Earlier work this paper cites.
Document ranking with a pretrained sequence-to-sequence model
Nogueira, R., Jiang, Z., Pradeep, R., and Lin, J · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Leveraging passage retrieval with generative models for open domain question answering
Izacard, G. and Grave, E · 2021
Earlier work this paper cites.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Jin, D., Pan, E., Oufattole, N., Weng, W.-H., Fang, H., and Szolovits, P · 2021
Earlier work this paper cites.
Sparse, dense, and attentional representations for text retrieval
Luan, Y., Eisenstein, J., Toutanova, K., and Collins, M · 2021
Earlier work this paper cites.
KILT: a benchmark for knowledge intensive language tasks
Petroni, F., Piktus, A., Fan, A., Lewis, P., Yazdani, M., De Cao, N., Thorne, J., Jernite, Y., Karpukhin, V., Maillard, J., Plachouras, V., Rocktäschel, T., and Riedel, S · 2021
Earlier work this paper cites.
End-to-end training of multi-document reader and retriever for open-domain question answering
Sachan, D. S., Reddy, S., Hamilton, W. L., Dyer, C., and Yogatama, D · 2021
Earlier work this paper cites.
Beir: A heterogeneous benchmark for zero-shot evaluation of information retrieval models
Thakur, N., Reimers, N., Rücklé, A., Srivastava, A., and Gurevych, I · 2021
Cited alongside, same era.
Tat-qa: A question answering benchmark on a hybrid of tabular and textual content in finance
Zhu, F., Lei, W., Huang, Y., Wang, C., Zhang, S., Lv, J., Feng, F., and Chua, T.-S · 2021
Cited alongside, same era.
Topiocqa: Open-domain conversational question answering with topic switching
Adlakha, V., Dhuliawala, S., Suleman, K., de Vries, H., and Reddy, S · 2022
Cited alongside, same era.
Improving language models by retrieving from trillions of tokens
Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Millican, K., Van Den Driessche, G. B., Lespiau, J.-B., Damoc, B., Clark, A., et al · 2022
Cited alongside, same era.
Glam: Efficient scaling of language models with mixture-of-experts
Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., et al · 2022
Cited alongside, same era.
Sail: Search-augmented instruction learning
Luo, H., Chuang, Y.-S., Gong, Y., Zhang, T., Kim, Y., Wu, X., Fox, D., Meng, H., and Glass, J · 2023
Later among the works it cites.
Fine-tuning llama for multi-stage text retrieval
Ma, X., Wang, L., Yang, N., Wei, F., and Lin, J · 2023
Later among the works it cites.
When not to trust language models: Investigating effectiveness of parametric and non-parametric memories
Mallen, A., Asai, A., Zhong, V., Das, R., Khashabi, D., and Hajishirzi, H · 2023
Later among the works it cites.
GPT-4, 2023
OpenAI · 2023
Later among the works it cites.
In-context retrieval-augmented language models
Ram, O., Levine, Y., Dalmedigos, I., Muhlgay, D., Shashua, A., Leyton-Brown, K., and Shoham, Y · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Re2G: Retrieve, rerank, generate
Glass, M., Rossiello, G., Chowdhury, M. F. M., Naik, A., Cai, P., and Gliozzo, A · 2022
Cited alongside, same era.
Unsupervised dense information retrieval with contrastive learning
Izacard, G., Caron, M., Hosseini, L., Riedel, S., Bojanowski, P., Joulin, A., and Grave, E · 2022
Cited alongside, same era.
Demonstrate-search-predict: Composing retrieval and language models for knowledge-intensive nlp
Khattab, O., Santhanam, K., Li, X. L., Hall, D., Liang, P., Potts, C., and Zaharia, M · 2022
Cited alongside, same era.
Internet-augmented language models through few-shot prompting for open-domain question answering
Lazaridou, A., Gribovskaya, E., Stokowiec, W., and Grigorev, N · 2022
Cited alongside, same era.
In defense of dual-encoders for neural ranking
Menon, A., Jayasumana, S., Rawat, A. S., Kim, S., Reddi, S., and Kumar, S · 2022
Cited alongside, same era.
Introducing ChatGPT, 2022
OpenAI · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Enhancing retrieval-augmented large language models with iterative retrieval-generation synergy
Shao, Z., Gong, Y., Shen, Y., Huang, M., Duan, N., and Chen, W · 2023
Later among the works it cites.
Is ChatGPT good at search? investigating large language models as re-ranking agents
Sun, W., Yan, L., Ma, X., Wang, S., Ren, P., Chen, Z., Yin, D., and Ren, Z · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions
Trivedi, H., Balasubramanian, N., Khot, T., and Sabharwal, A · 2023
Later among the works it cites.
Inscit: Information-seeking conversations with mixed-initiative interactions
Wu, Z., Parish, R., Cheng, H., Min, S., Ammanabrolu, P., Ostendorf, M., and Hajishirzi, H · 2023
Later among the works it cites.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al · 2024
Closest in time.
Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024
DeepSeek · 2024
Closest in time.
Adaptive-rag: Learning to adapt retrieval-augmented large language models through question complexity
Jeong, S., Baek, J., Cho, S., Hwang, S. J., and Park, J. C · 2024
Closest in time.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., et al · 2024
Closest in time.
Nv-embed: Improved techniques for training llms as generalist embedding models
Lee, C., Roy, R., Xu, M., Raiman, J., Shoeybi, M., Catanzaro, B., and Ping, W · 2024
Closest in time.
RA-DIT: Retrieval-augmented dual instruction tuning
Lin, X. V., Chen, X., Chen, M., Shi, W., Lomeli, M., James, R., Rodriguez, P., Kahn, J., Szilvasy, G., Lewis, M., Zettlemoyer, L., and tau Yih, W · 2024
Closest in time.
Chatqa: Surpassing gpt-4 on conversational qa and rag
Liu, Z., Ping, W., Roy, R., Xu, P., Shoeybi, M., and Catanzaro, B · 2024
Closest in time.
Llama 3 model card
Meta-AI · 2024
Closest in time.
Mixtral 8x22b
Mistral · 2024
Closest in time.
Generative representational instruction tuning
Muennighoff, N., Su, H., Wang, L., Yang, N., Wei, F., Yu, T., Singh, A., and Kiela, D · 2024
Closest in time.
Proving test set contamination in black-box language models
Oren, Y., Meister, N., Chatterji, N. S., Ladhak, F., and Hashimoto, T · 2024
Closest in time.
Large language models are effective text rankers with pairwise ranking prompting
Qin, Z., Jagerman, R., Hui, K., Zhuang, H., Wu, J., Shen, J., Liu, T., Liu, J., Metzler, D., Wang, X., et al · 2024
Closest in time.
Replug: Retrieval-augmented black-box language models
Shi, W., Min, S., Yasunaga, M., Seo, M., James, R., Lewis, M., Zettlemoyer, L., and Yih, W.-t · 2024
Closest in time.
Instructretro: Instruction tuning post retrieval-augmented pretraining
Wang, B., Ping, W., McAfee, L., Xu, P., Li, B., Shoeybi, M., and Catanzaro, B · 2024
Closest in time.
Pmc-llama: toward building open-source language models for medicine
Wu, C., Lin, W., Zhang, X., Zhang, Y., Xie, W., and Wang, Y · 2024
Closest in time.
Benchmarking retrieval-augmented generation for medicine
Xiong, G., Jin, Q., Lu, Z., and Zhang, A · 2024
Closest in time.
Making retrieval-augmented language models robust to irrelevant context
Yoran, O., Wolfson, T., Ram, O., and Berant, J · 2024
Closest in time.
Improving language models via plug-and-play retrieval feedback, 2024
Yu, W., Zhang, Z., Liang, Z., Jiang, M., and Sabharwal, A · 2024
Closest in time.
Raft: Adapting language model to domain specific rag
Zhang, T., Patil, S. G., Jain, N., Shen, S., Zaharia, M., Stoica, I., and Gonzalez, J. E · 2024
Closest in time.
Inters: Unlocking the power of large language models in search with instruction tuning
Zhu, Y., Zhang, P., Zhang, C., Chen, Y., Xie, B., Dou, Z., Liu, Z., and Wen, J.-R · 2024
Closest in time.