Fetching the paper…
Reading the bibliography…
Existing RAG benchmarks often overlook query difficulty, leading to inflated performance on simpler questions and unreliable evaluations.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, and 1 others. 2020 · 1901
Earlier work this paper cites.
Kilt: a benchmark for knowledge intensive language tasks
Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, and 1 others. 2020 · 2009
Earlier work this paper cites.
Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps
Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. 2020 · 2011
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018 · 2018
Earlier work this paper cites.
Question answering by reasoning across documents with graph convolutional networks
Nicola De Cao, Wilker Aziz, and Ivan Titov. 2019 · 2019
Earlier work this paper cites.
Answering complex questions by joining multi-document evidence with quasi knowledge graphs
Xiaolu Lu, Soumajit Pramanik, Rishiraj Saha Roy, Abdalghani Abujabal, Yafang Wang, and Gerhard Weikum. 2019 · 2019
Earlier work this paper cites.
Beat the ai: Investigating adversarial human annotation for reading comprehension
Max Bartolo, Alastair Roberts, Johannes Welbl, Sebastian Riedel, and Pontus Stenetorp. 2020 · 2020
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020 · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, and 1 others. 2020 · 2020
Earlier work this paper cites.
Leveraging passage retrieval with generative models for open domain question answering
Gautier Izacard and Édouard Grave. 2021 · 2021
Earlier work this paper cites.
Re2g: Retrieve, rerank, generate
Michael Glass, Gaetano Rossiello, Md Faisal Mahbub Chowdhury, Ankita Naik, Pengshan Cai, and Alfio Gliozzo. 2022 · 2022
Cited alongside, same era.
Musique: Multihop questions via single-hop question composition
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022 · 2022
Cited alongside, same era.
Retrieval-augmented generation for large language models: A survey
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Haofen Wang, and Haofen Wang. 2023 · 2023
Cited alongside, same era.
Factkg: Fact verification via reasoning on knowledge graphs
Jiho Kim, Sungjin Park, Yeonsu Kwon, Yohan Jo, James Thorne, and Edward Choi. 2023 · 2023
Cited alongside, same era.
Ares: An automated evaluation framework for retrieval-augmented generation systems
Jon Saad-Falcon, Omar Khattab, Christopher Potts, and Matei Zaharia. 2023 · 2023
Raptor: Recursive abstractive processing for tree-organized retrieval
Parth Sarthi, Salman Abdullah, Aditi Tuli, Shubh Khanna, Anna Goldie, and Christopher D Manning. 2024 · 2024
Later among the works it cites.
An interactive question answer based system on alzheimer’s disease using retrieval augmented generation
Sujoy Sen, Samay Sarkar, Partha Ghosh, Takaaki Goto, and Soumya Sen. 2024 · 2024
Later among the works it cites.
Retrieval augmented generation for domain-specific question answering
Sanat Sharma, David Seunghyun Yoon, Franck Dernoncourt, Dewang Sultania, Karishma Bagga, Mengjiao Zhang, Trung Bui, and Varun Kotte. 2024 · 2024
Later among the works it cites.
A methodology for evaluating rag systems: A case study on configuration dependency validation
Sebastian Simon, Alina Mailach, Johannes Dorn, and Norbert Siegmund. 2024 · 2024
Later among the works it cites.
Multihop-rag: Benchmarking retrieval-augmented generation for multi-hop queries
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Performance prediction for multi-hop questions
Mohammadreza Samadi and Davood Rafiei. 2023 · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, and 1 others. 2023 · 2023
Cited alongside, same era.
Conqret: Benchmarking fine-grained evaluation of retrieval augmented argumentation with llm judges
Kaustubh D Dhole, Kai Shu, and Eugene Agichtein. 2024 · 2024
Cited alongside, same era.
From local to global: A graph rag approach to query-focused summarization
Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, Dasha Metropolitansky, Robert Osazuwa Ness, and Jonathan Larson. 2024 · 2024
Cited alongside, same era.
l o n g 2 r a g long^{2}rag : Evaluating long-context & long-form retrieval-augmented generation with key point recall
Zehan Qi, Rongwu Xu, Zhijiang Guo, Cunxiang Wang, Hao Zhang, and Wei Xu. 2024 · 2024
Cited alongside, same era.
Yixuan Tang and Yi Yang. 2024 · 2024
Later among the works it cites.
Evaluation of retrieval-augmented generation: A survey
Hao Yu, Aoran Gan, Kai Zhang, Shiwei Tong, Qi Liu, and Zhaofeng Liu. 2024 · 2024
Later among the works it cites.
Shuliang Liu, Xinze Li, Zhenghao Liu, Yukun Yan, Cheng Yang, Zheni Zeng, Zhiyuan Liu, Maosong Sun, and Ge Yu. 2025 · 2025
Closest in time.
Crud-rag: A comprehensive chinese benchmark for retrieval-augmented generation of large language models
Yuanjie Lyu, Zhiyu Li, Simin Niu, Feiyu Xiong, Bo Tang, Wenjin Wang, Hao Wu, Huanyong Liu, Tong Xu, and Enhong Chen. 2025 · 2025
Closest in time.
Using the retrieval-augmented generation to improve the question-answering system in human health risk assessment: The development and application
Wenjun Meng, Yuzhe Li, Lili Chen, and Zhaomin Dong. 2025 · 2025
Closest in time.
Ragchecker: A fine-grained framework for diagnosing retrieval-augmented generation
Dongyu Ru, Lin Qiu, Xiangkun Hu, Tianhang Zhang, Peng Shi, Shuaichen Chang, Cheng Jiayang, Cunxiang Wang, Shichao Sun, Huanyu Li, and 1 others. 2025 · 2025
Closest in time.