Fetching the paper…
Reading the bibliography…
We present GaRAGe, a large RAG benchmark with human-curated long-form answers and annotations of each grounding passage, allowing a fine-grained evaluation of whether LLMs can identify relevant grounding when generating RAG answers.
Introducing the enron corpus
Bryan Klimt and Yiming Yang. 2004 · 2004
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
Wikiqa: A challenge dataset for open-domain question answering
Yi Yang, Wen-tau Yih, and Christopher Meek. 2015 · 2018
Earlier work this paper cites.
Eli5: Long form question answering
Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019 · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019 · 2019
Earlier work this paper cites.
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020 · 2020
Earlier work this paper cites.
SituatedQA: Incorporating extra-linguistic contexts into QA
Michael J.Q. Zhang and Eunsol Choi. 2021 · 2021
Earlier work this paper cites.
MuSiQue: Multihop questions via single-hop question composition
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022 · 2022
Earlier work this paper cites.
Reasoning over public and private data in retrieval-based systems
Simran Arora, Patrick Lewis, Angela Fan, Jacob Kahn, and Christopher Ré. 2023 · 2023
Earlier work this paper cites.
Robustqa: Benchmarking the robustness of domain adaptation for open-domain question answering
Rujun Han, Peng Qi, Yuhao Zhang, Lan Liu, Juliette Burger, William Yang Wang, Zhiheng Huang, Bing Xiang, and Dan Roth. 2023 · 2023
Earlier work this paper cites.
Recall: A benchmark for llms robustness against external counterfactual knowledge
Yi Liu, Lianzhe Huang, Shicheng Li, Sishuo Chen, Hao Zhou, Fandong Meng, Jie Zhou, and Xu Sun. 2023 · 2023
Cited alongside, same era.
Detrimental contexts in open-domain question answering
Philhoon Oh and James Thorne. 2023 · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al. 2023 · 2023
Cited alongside, same era.
Introducing the next generation of claude
AI Anthropic. 2024 · 2024
Cited alongside, same era.
Ragas: Automated evaluation of retrieval augmented generation
Shahul Es, Jithin James, Luis Espinosa Anke, and Steven Schockaert. 2024 · 2024
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024 · 2024
Later among the works it cites.
Realtime qa: what’s the answer right now?
Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Akari Asai, Xinyan Yu, Dragomir Radev, Noah A Smith, Yejin Choi, Kentaro Inui, et al. 2024 · 2024
Later among the works it cites.
Summary of a haystack: A challenge to long-context LLMs and RAG systems
Philippe Laban, Alexander Fabbri, Caiming Xiong, and Chien-Sheng Wu. 2024 · 2024
Later among the works it cites.
Llms as narcissistic evaluators: When ego inflates evaluation scores
Yiqi Liu, Nafise Moosavi, and Chenghua Lin. 2024 · 2024
Later among the works it cites.
RAGTruth: A hallucination corpus for developing trustworthy retrieval-augmented language models
Cheng Niu, Yuanhao Wu, Juno Zhu, Siliang Xu, KaShun Shum, Randy Zhong, Juntong Song, and Tong Zhang. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Measuring retrieval complexity in question answering systems
Matteo Gabburo, Nicolaas Paul Jedema, Siddhant Garg, Leonardo F. R. Ribeiro, and Alessandro Moschitti. 2024 · 2024
Cited alongside, same era.
Automated evaluation of retrieval-augmented language models with task-specific exam generation
Gauthier Guinet, Behrooz Omidvar-Tehrani, Anoop Deoras, and Laurent Callot. 2024 · 2024
Cited alongside, same era.
Rag-qa arena: Evaluating domain robustness for long-form retrieval augmented question answering
Rujun Han, Yuhao Zhang, Peng Qi, Yumo Xu, Jenyuan Wang, Lan Liu, William Yang Wang, Bonan Min, and Vittorio Castelli. 2024 · 2024
Cited alongside, same era.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. 2024 · 2024
Cited alongside, same era.
The amazon nova family of models: Technical report and model card
Amazon Artificial General Intelligence. 2024 · 2024
Cited alongside, same era.
FACTS grounding: A new benchmark for evaluating the factuality of large language models
Alon Jacovi, Andrew Wang, Chris Alberti, Connie Tao, Jon Lipovetz, Kate Olszewska, Lukas Haas, Michelle Liu, Nate Keating, Adam Bloniarz, Carl Saroufim, Corey Fry, Dror Marcus, Doron Kukliansky, Gaurav Singh Tomar, James Swirhun, Jinwei Xing, Lily Wang, Madhu Gurumurthy, Michael Aaron, Moran Ambar, Rachana Fellinger, Rui Wang, Zizhao Zhang, Sasha Goldshtein, and Dipanjan Das. 2024 · 2024
Cited alongside, same era.
Jianlv Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, and Zheng Liu. 2024a
Cited in the paper.
Ronak Pradeep, Nandan Thakur, Sahel Sharifymoghaddam, Eric Zhang, Ryan Nguyen, Daniel Campos, Nick Craswell, and Jimmy Lin. 2024 · 2024
Later among the works it cites.
Ares: An automated evaluation framework for retrieval-augmented generation systems
Jon Saad-Falcon, Omar Khattab, Christopher Potts, and Matei Zaharia. 2024 · 2024
Later among the works it cites.
Multihop-RAG: Benchmarking retrieval-augmented generation for multi-hop queries
Yixuan Tang and Yi Yang. 2024 · 2024
Later among the works it cites.
Michelangelo: Long context evaluations beyond haystacks via latent structure queries
Kiran Vodrahalli, Santiago Ontanon, Nilesh Tripuraneni, Kelvin Xu, Sanil Jain, Rakesh Shivanna, Jeffrey Hui, Nishanth Dikkala, Mehran Kazemi, Bahare Fatemi, et al. 2024 · 2024
Later among the works it cites.
Kbqa-o1: Agentic knowledge base question answering with monte carlo tree search
Haoran Luo, Haihong E, Yikai Guo, Qika Lin, Xiaobao Wu, Xinyu Mu, Wenhao Liu, Meina Song, Yifan Zhu, and Luu Anh Tuan. 2025 · 2025
Closest in time.