Fetching the paper…
Reading the bibliography…
Retrieval-Augmented Generation (RAG) significantly improves document-based question answering by integrating external documents during generation.
Voorhees, E.M., et al
1999
Earlier work this paper cites.
Järvelin, K., Kekäläinen, J.: Cumulated gain-based evaluation of ir techniques. ACM Transactions on Information Systems (TOIS) 20
2002
Earlier work this paper cites.
Schütze, H., Manning, C.D., Raghavan, P.: Introduction to Information Retrieval vol. 39. Cambridge University Press Cambridge, Cambridge, UK (2008)
2008
Earlier work this paper cites.
Rajpurkar, P., Zhang, J., Lopyrev, K., Liang, P.: Squad: 100,000+ questions for machine comprehension of text. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2383–2392 (2016) https://doi.org/10.18653/v1/D16-1264
2016
Earlier work this paper cites.
Nguyen, T., Rosenberg, M., Song, X., Gao, J., Tiwary, S., Majumder, R., Deng, L.: Ms marco: A human generated machine reading comprehension dataset. In: Proceedings of the Workshop on Cognitive Computation: Integrating Neural and Symbolic Approaches at NIPS 2016 (2016)
2016
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. In: Advances in Neural Information Processing Systems (NeurIPS), vol. 30, pp. 5998–6008 (2017)
2017
Earlier work this paper cites.
Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W.W., Salakhutdinov, R., Manning, C.D.: HotpotQA: A dataset for diverse, explainable multi-hop question answering. In: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP) (2018). https://doi.org/10.18653/v1/D18-1259
2018
Earlier work this paper cites.
Kočiský, T., Schwarz, J., Blunsom, P., Dyer, C., Hermann, K.M., Melis, G., Grefenstette, E.: The narrativeqa reading comprehension challenge. Transactions of the Association for Computational Linguistics 6
2018
Earlier work this paper cites.
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., et al
2019
Earlier work this paper cites.
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al
2020
Earlier work this paper cites.
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., Riedel, S., Kiela, D.: Retrieval-augmented generation for knowledge-intensive nlp tasks. In: Advances in Neural Information Processing Systems (NeurIPS), vol. 33, pp. 9459–9474 (2020)
2020
Earlier work this paper cites.
Xiong, W., Li, X.L., Iyer, S., Du, J., Lewis, P., Wang, W.Y., Mehdad, Y., Yih, W.-t., Riedel, S., Kiela, D., et al
2021
Earlier work this paper cites.
Dasigi, P., Liu, N.F., Peters, M.E., Gardner, M.: A dataset of information-seeking questions and answers anchored in research papers. In: Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL), pp. 4599–4610 (2021)
2021
Earlier work this paper cites.
Glass, M., Rossiello, G., Chowdhury, M.F.M., Naik, A.R., Cai, P., Gliozzo, A.: Re2g: Retrieve, rerank, generate. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL), pp. 2701–2715 (2022). https://doi.org/10.18653/v1/2022.naacl-main.194
2022
Earlier work this paper cites.
Ye, M., Shen, J., Lin, G., Xiang, T., Shao, L., Hoi, S.C.H.: Bi-level inter-modality modulation for unsupervised visible-infrared person re-identification. Pattern Recognition 131
2022
Earlier work this paper cites.
Liu, J.: Llamaindex (2022) https://doi.org/10.5281/zenodo.1234
2022
Earlier work this paper cites.
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H.W., Sutton, C., Gehrmann, S., et al
2023
Cited alongside, same era.
Lála, J., O’Donoghue, O., Shtedritski, A., Cox, S., Rodriques, S.G., White, A.D.: PaperQA: Retrieval-Augmented Generative Agent for Scientific Research (2023). https://doi.org/10.48550/arXiv.2312.07559
2023
Cited alongside, same era.
Rajabzadeh, H., Wang, S., Kwon, H.J., Liu, B.: Multimodal Multi-Hop Question Answering Through a Conversation Between Tools and Efficiently Finetuned Large Language Models (2023)
2023
Cited alongside, same era.
Shi, F., Chen, X., Misra, K., Scales, N., Dohan, D., Chi, E., Schärli, N., Zhou, D.: Large language models can be easily distracted by irrelevant context. In: Proceedings of the 40th International Conference on Machine Learning (ICML), vol. 202, pp. 31210–31227 (2023)
2023
Cited alongside, same era.
Rackauckas, D.: Rag-fusion: A new take on retrieval-augmented generation. International Journal on Natural Language Computing (IJNLC) 13
2024
Closest in time.
Liu, N.F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., Liang, P.: Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics (TACL) 12
2024
Closest in time.
Yan, S.-Q., Gu, J.-C., Zhu, Y., Ling, Z.-H.: Corrective retrieval augmented generation. In: Findings of the Association for Computational Linguistics: EMNLP 2024 (2024). https://doi.org/10.18653/V1/2024.FINDINGS-EMNLP.123
2024
Closest in time.
Xu, Z., Liu, Z., Yan, Y., Wang, S., Yu, S., Zeng, Z., Xiao, C., Liu, Z., Yu, G., Xiong, C.: ActiveRAG: Autonomously Knowledge Assimilation and Accommodation through Retrieval-Augmented Agents (2024). https://doi.org/10.48550/arXiv.2402.13547
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jiang, Z., Xu, F.F., Gao, L., Sun, Z., Liu, Q., Dwivedi-Yu, J., Yang, Y., Callan, J., Neubig, G.: Active retrieval augmented generation. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 7969–7992 (2023). https://doi.org/10.18653/v1/2023.emnlp-main.495
2023
Cited alongside, same era.
Li, X., Zhang, W., Chen, J.: Hierarchical sequential context modelling for high-fidelity image inpainting. Image and Vision Computing 136
2023
Cited alongside, same era.
Zhao, B., Ji, C., Zhang, Y., He, W., Wang, Y., Wang, Q., Feng, R., Zhang, X.: Large language models are complex table parsers. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 14786–14802 (2023). https://doi.org/10.18653/v1/2023.emnlp-main.914
2023
Cited alongside, same era.
Bruch, S., Gai, S., Ingber, A.: An analysis of fusion functions for hybrid retrieval. ACM Transactions on Information Systems 42
2023
Cited alongside, same era.
Saad-Falcon, J., Barrow, J., Siu, A., Nenkova, A., Yoon, D.S., Rossi, R.A., Dernoncourt, F.: Pdftriage: Question answering over long, structured documents. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track, pp. 153–169 (2024). https://doi.org/10.18653/v1/2024.emnlp-industry.13
2024
Cited alongside, same era.
OpenAI: Introducing GPT-4o. Accessed: 2025-03-14 (2024). https://openai.com/blog/introducing-gpt-4o
2024
Cited alongside, same era.
Cuconasu, F., Trappolini, G., Siciliano, F., Filice, S., Campagnano, C., Maarek, Y., Tonellotto, N., Silvestri, F.: The power of noise: Redefining retrieval for rag systems. In: Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 719–729 (2024). https://doi.org/10.1145/3626772.3657834
2024
Cited alongside, same era.
Yoran, O., Wolfson, T., Ram, O., Berant, J.: Making retrieval-augmented language models robust to irrelevant context. In: The Twelfth International Conference on Learning Representations (ICLR) (2024)
2024
Cited alongside, same era.
Sarthi, P., Abdullah, S., Tuli, A., Khanna, S., Goldie, A., Manning, C.D.: Raptor: Recursive abstractive processing for tree-organized retrieval. In: The Twelfth International Conference on Learning Representations (ICLR) (2024)
2024
Closest in time.
Günther, M., Mohr, I., Williams, D.J., Wang, B., Xiao, H.: Late chunking: Contextual chunk embeddings using long-context embedding models. In: Findings of the Association for Computational Linguistics: EMNLP 2024 (2024)
2024
Closest in time.
Pu, N., Chen, W., Liu, Y., Bakker, E.M., Lew, M.S.: Unsupervised lifelong person re-identification via affinity harmonization. International Journal of Computer Vision 132
2024
Closest in time.
Wang, B., Xu, C., Zhao, X., Ouyang, L., Wu, F., Zhao, Z., Xu, R., Liu, K., Qu, Y., Shang, F., Zhang, B., Wei, L., Sui, Z., Li, W., Shi, B., Qiao, Y., Lin, D., He, C.: MinerU: An Open-Source Solution for Precise Document Content Extraction (2024). https://doi.org/10.48550/arXiv.2409.18839
2024
Closest in time.
Elastic: Accelerate time to insight with Elasticsearch and AI. https://www.elastic.co/ (2024)
2024
Closest in time.
Hsu, W., Tzeng, E.: Dynamic Alpha Tuning: Enhancing Hybrid Retrieval with Adaptive Weighting (2025)
2025
Closest in time.
Merola, C., Singh, J.: Reconstructing context: Evaluating advanced chunking strategies for retrieval-augmented generation. In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025) (2025)
2025
Closest in time.
Chen, H., Yang, Y., Li, Y., Zhang, M., Hu, B., Zhang, M.: Beyond Chunking: Discourse-Aware Hierarchical Retrieval for Long Document Question Answering (2025). https://doi.org/10.48550/arXiv.2506.06313
2025
Closest in time.
Gong, Z., Mai, C., Huang, Y.: MHier-RAG: Multi-Modal RAG for Visual-Rich Document Question-Answering via Hierarchical and Multi-Granularity Reasoning (2025)
2025
Closest in time.
ChatPDF: ChatPDF: AI-Powered PDF Interaction Tool (2025). https://www.chatpdf.com/
2025
Closest in time.
Sun, L., Zhang, L., Huang, J., Jia, C., Cheng, Z., Zhang, X.: FABLE: Forest-Based Adaptive Bi-Path LLM-Enhanced Retrieval for Multi-Document Reasoning (2026)
2026
Closest in time.