Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have significantly enhanced the capabilities of information access systems, especially with retrieval-augmented generation (RAG).
On the Evaluation of Machine-Generated Reports. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2024) . Washington, D.C., 1904–1915
James Mayfield, Eugene Yang, Dawn Lawrie, Sean MacAvaney, Paul McNamee, Douglas W. Oard, Luca Soldaini, Ian Soboroff, Orion Weller, Efsun Kayi, Kate Sanders, Marc Mason, and Noah Hibbler. 2024 · 1915
Earlier work this paper cites.
Variations in Relevance Judgments and the Measurement of Retrieval Effectiveness. In Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 1998) . Melbourne, Australia, 315–323
Ellen M. Voorhees. 1998 · 1998
Earlier work this paper cites.
How Reliable Are the Results of Large-Scale Information Retrieval Experiments?. In Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 1998) . Melbourne, Australia, 307–314
Justin Zobel. 1998 · 1998
Earlier work this paper cites.
Evaluating Evaluation Measure Stability. In Proceedings of the 23rd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2000) . Athens, Greece, 33–40
Chris Buckley and Ellen M. Voorhees. 2000 · 2000
Earlier work this paper cites.
Evaluating Answers to Definition Questions. In Companion Volume of the Proceedings of HLT-NAACL 2003 — Short Papers . Edmonton, Canada, 109–111
Ellen M. Voorhees. 2003a · 2003
Earlier work this paper cites.
Overview of the TREC 2003 Question Answering Track. In Proceedings of the Twelfth Text REtrieval Conference (TREC 2003) . Gaithersburg, Maryland
Ellen M. Voorhees. 2003b · 2003
Earlier work this paper cites.
Retrieval Evaluation with Incomplete Information. In Proceedings of the 27th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2004) . Sheffield, United Kingdom, 25–32
Chris Buckley and Ellen M. Voorhees. 2004 · 2004
Earlier work this paper cites.
Evaluating Content Selection in Summarization: The Pyramid Method. In Proceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics: HLT-NAACL 2004 . Boston, Massachusetts, 145–152
Ani Nenkova and Rebecca Passonneau. 2004 · 2004
Earlier work this paper cites.
Automatically Evaluating Answers to Definition Questions. In Proceedings of the 2005 Human Language Technology Conference and Conference on Empirical Methods in Natural Language Processing (HLT/EMNLP 2005) . Vancouver, Canada, 931–938
Jimmy Lin and Dina Demner-Fushman. 2005 · 2005
Earlier work this paper cites.
Information Retrieval System Evaluation: Effort, Sensitivity, and Reliability. In Proceedings of the 28th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2005) . Salvador, Brazil, 162–169
Mark Sanderson and Justin Zobel. 2005 · 2005
Earlier work this paper cites.
Methods for Automatically Evaluating Answers to Complex Questions
Jimmy Lin and Dina Demner-Fushman. 2006a · 2006
Earlier work this paper cites.
Different Structures for Evaluating Answers to Complex Questions: Pyramids Won’t Topple, and Neither Will Human Assessors. In Proceedings of the 45th Annual Meeting of the Association for Computational Linguistics (ACL 2007) . Prague, Czech Republic, 768–775
Hoa Trang Dang and Jimmy Lin. 2007 · 2007
Earlier work this paper cites.
Deconstructing Nuggets: The Stability and Reliability of Complex Question Answering Evaluation. In Proceedings of the 30th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2007) . Amsterdam, the Netherlands, 327–334
Jimmy Lin and Pengyi Zhang. 2007 · 2007
Earlier work this paper cites.
Information Retrieval Evaluation
Donna Harman. 2011 · 2011
Earlier work this paper cites.
IR System Evaluation using Nugget-based Test Collections. In Proceedings of the Fifth ACM International Conference on Web Search and Data Mining (WSDM 2012) . Seattle, Washington, 393–402
Virgil Pavlu, Shahzad Rajput, Peter B. Golbus, and Javed A. Aslam. 2012 · 2012
Earlier work this paper cites.
REALM: Retrieval-Augmented Language Model Pre-training. In Proceedings of the 37th International Conference on Machine Learning (ICML 2020) . 3929–3938
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020 · 2020
Earlier work this paper cites.
Generalization through Memorization: Nearest Neighbor Language Models. In Proceedings of the 8th International Conference on Learning Representations (ICLR 2020) . Addis Ababa, Ethiopia
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020 · 2020
Earlier work this paper cites.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Advances in Neural Information Processing Systems 33 . 9459–9474
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020 · 2020
Cited alongside, same era.
Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume . 874–880
Gautier Izacard and Edouard Grave. 2021 · 2021
Cited alongside, same era.
Improving Language Models by Retrieving from Trillions of Tokens. In Proceedings of the 39th International Conference on Machine Learning (ICML 2022) . Baltimore, Maryland, 2206–2240
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, et al · 2022
Cited alongside, same era.
Enabling Large Language Models to Generate Text with Citations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . Singapore, 6465–6488
Tianyu Gao, Howard Yen, Jiatong Yu, and Danqi Chen. 2023 · 2023
SFR-RAG: Towards Contextually Faithful LLMs
Xuan-Phi Nguyen, Shrey Pandit, Senthil Purushwalkam, Austin Xu, Hailin Chen, Yifei Ming, Zixuan Ke, Silvio Savarese, Caiming Xong, and Shafiq Joty. 2024 · 2024
Later among the works it cites.
ConvKGYarn: Spinning Configurable and Scalable Conversational Knowledge Graph QA Datasets with Large Language Models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track . Miami, Florida, 1176–1206
Ronak Pradeep, Daniel Lee, Ali Mousavi, Jeffrey Pound, Yisi Sang, Jimmy Lin, Ihab Ilyas, Saloni Potdar, Mostafa Arefiyan, and Yunyao Li. 2024a · 2024
Later among the works it cites.
Initial Nugget Evaluation Results for the TREC 2024 RAG Track with the AutoNuggetizer Framework
Ronak Pradeep, Nandan Thakur, Shivani Upadhyay, Daniel Campos, Nick Craswell, and Jimmy Lin. 2024b · 2024
Later among the works it cites.
Evaluating RAG-Fusion with RAGElo: an Automated Elo-based Framework
Zackary Rackauckas, Arthur Câmara, and Jakub Zavrel. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . Singapore, 12076–12100
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023 · 2023
Cited alongside, same era.
A Critical Evaluation of Evaluations for Long-form Question Answering. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Toronto, Canada, 3225–3245
Fangyuan Xu, Yixiao Song, Mohit Iyyer, and Eunsol Choi. 2023 · 2023
Cited alongside, same era.
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. In Advances in Neural Information Processing Systems 36 (NeurIPS 2023) Datasets and Benchmarks Track . New Orleans, Louisiana, 46595–46623
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2023
Cited alongside, same era.
Generative information retrieval evaluation
Marwah Alaofi, Negar Arabzadeh, Charles LA Clarke, and Mark Sanderson. 2024 · 2024
Cited alongside, same era.
A Comparison of Methods for Evaluating Generative IR
Negar Arabzadeh and Charles LA Clarke. 2024 · 2024
Cited alongside, same era.
Benchmarking Large Language Models in Retrieval-Augmented Generation. In Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence (AAAI-24) . Vancouver, Canada, 17754–17762
Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun. 2024 · 2024
Cited alongside, same era.
RAGAs: Automated Evaluation of Retrieval Augmented Generation. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations . St. Julians, Malta, 150–158
Shahul Es, Jithin James, Luis Espinosa Anke, and Steven Schockaert. 2024 · 2024
Cited alongside, same era.
Pencils Down! Automatic Rubric-based Evaluation of Retrieve/Generate Systems. In Proceedings of the 2024 ACM SIGIR International Conference on Theory of Information Retrieval (ICTIR ’24) . Washington, D.C., 175–184
Naghmeh Farzi and Laura Dietz. 2024 · 2024
Cited alongside, same era.
Later among the works it cites.
Researchy Questions: A Dataset of Multi-Perspective, Decompositional Questions for LLM Web Agents
Corby Rosset, Ho-Lam Chung, Guanghui Qin, Ethan C. Chau, Zhuo Feng, Ahmed Awadallah, Jennifer Neville, and Nikhil Rao. 2024 · 2024
Later among the works it cites.
ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) . Mexico City, Mexico, 338–354
Jon Saad-Falcon, Omar Khattab, Christopher Potts, and Matei Zaharia. 2024 · 2024
Later among the works it cites.
Evaluating Retrieval Quality in Retrieval-Augmented Generation
Alireza Salemi and Hamed Zamani. 2024 · 2024
Later among the works it cites.
RAGProbe: An Automated Approach for Evaluating RAG Applications
Shangeetha Sivasothy, Scott Barnett, Stefanus Kurniawan, Zafaryab Rasool, and Rajesh Vasa. 2024 · 2024
Later among the works it cites.
“Knowing When You Don‘t Know”: A Multilingual Relevance Assessment Dataset for Robust Retrieval-Augmented Generation. In Findings of the Association for Computational Linguistics: EMNLP 2024 . Miami, Florida, 12508–12526
Nandan Thakur, Luiz Bonifacio, Crystina Zhang, Odunayo Ogundepo, Ehsan Kamalloo, David Alfonso-Hermelo, Xiaoguang Li, Qun Liu, Boxing Chen, Mehdi Rezagholizadeh, and Jimmy Lin. 2024 · 2024
Later among the works it cites.
A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look
Shivani Upadhyay, Ronak Pradeep, Nandan Thakur, Daniel Campos, Nick Craswell, Ian Soboroff, Hoa Trang Dang, and Jimmy Lin. 2024a · 2024
Later among the works it cites.
UMBRELA: UMbrela is the (Open-Source Reproduction of the) Bing RELevance Assessor
Shivani Upadhyay, Ronak Pradeep, Nandan Thakur, Nick Craswell, and Jimmy Lin. 2024b · 2024
Later among the works it cites.
Evaluating Quality of Answers for Retrieval-Augmented Generation: A Strong LLM Is All You Need
Yang Wang, Alberto Garcia Hernandez, Roman Kyslyi, and Nicholas Kersting. 2024 · 2024
Later among the works it cites.
CRAG – Comprehensive RAG Benchmark
Xiao Yang, Kai Sun, Hao Xin, Yushi Sun, Nikita Bhalla, Xiangsen Chen, Sajal Choudhary, Rongze Daniel Gui, Ziran Will Jiang, Ziyu Jiang, Lingkun Kong, Brian Moran, Jiaqi Wang, Yifan Ethan Xu, An Yan, Chenyu Yang, Eting Yuan, Hanwen Zha, Nan Tang, Lei Chen, Nicolas Scheffer, Yue Liu, Nirav Shah, Rakesh Wanga, Anuj Kumar, Wen tau Yih, and Xin Luna Dong. 2024 · 2024
Later among the works it cites.
Ragnarök: A Reusable RAG Framework and Baselines for TREC 2024 Retrieval-Augmented Generation Track. In Proceedings of the 47th European Conference on Information Retrieval (ECIR 2025), Part I . Lucca, Italy, 132–148
Ronak Pradeep, Nandan Thakur, Sahel Sharifymoghaddam, Eric Zhang, Ryan Nguyen, Daniel Campos, Nick Craswell, and Jimmy Lin. 2025 · 2025
Closest in time.
MIRAGE-Bench: Automatic Multilingual Benchmark Arena for Retrieval-Augmented Generation Systems
Nandan Thakur, Suleman Kazi, Ge Luo, Jimmy Lin, and Amin Ahmad. 2025a · 2025
Closest in time.
FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents
Nandan Thakur, Jimmy Lin, Sam Havens, Michael Carbin, Omar Khattab, and Andrew Drozdov. 2025b · 2025
Closest in time.
CodeRAG-Bench: Can Retrieval Augment Code Generation?
Zora Zhiruo Wang, Akari Asai, Xinyan Velocity Yu, Frank F. Xu, Yiqing Xie, Graham Neubig, and Daniel Fried. 2025 · 2025
Closest in time.