Fetching the paper…
Reading the bibliography…
Challenges in the automated evaluation of Retrieval-Augmented Generation (RAG) Question-Answering (QA) systems include hallucination problems in domain-specific knowledge and the lack of gold standard benchmarks for company internal tasks.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019 · 1910
Earlier work this paper cites.
The treatment of ties in ranking problems
M G KENDALL. 1945 · 1945
Earlier work this paper cites.
Okapi at TREC-3. In Proceedings of The Third Text REtrieval Conference, TREC 1994, Gaithersburg, Maryland, USA, November 2-4, 1994 (NIST Special Publication, Vol. 500-225) , Donna K. Harman (Ed.). National Institute of Standards and Technology (NIST), 109–126
Stephen E. Robertson, Steve Walker, Susan Jones, Micheline Hancock-Beaulieu, and Mike Gatford. 1994 · 1994
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation. In ACL 2002 (Philadelphia, Pennsylvania) (ACL ’02) . Association for Computational Linguistics, USA, 311–318
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out . Association for Computational Linguistics, Barcelona, Spain, 74–81
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Language Models Are Few-Shot Learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2005
Earlier work this paper cites.
Meteor: an automatic metric for MT evaluation with high levels of correlation with human judgments. In StatMT 2007 (Prague, Czech Republic) (StatMT ’07) . Association for Computational Linguistics, USA, 228–231
Alon Lavie and Abhaya Agarwal. 2007 · 2007
Earlier work this paper cites.
Correlation Coefficients: Appropriate Use and Interpretation
Patrick Schober, Christa Boer, and Lothar A Schwarte. 2018 · 2018
Earlier work this paper cites.
Pairwise Crowd Judgments: Preference, Absolute, and Ratio. In Proceedings of the 23rd Australasian Document Computing Symposium (Dunedin, New Zealand) (ADCS ’18) . Association for Computing Machinery, New York, NY, USA, Article 3, 8 pages
Ziying Yang, Alistair Moffat, and Andrew Turpin. 2018 · 2018
Earlier work this paper cites.
Reciprocal rank fusion outperforms condorcet and individual rank learning methods. In SIGIR 2019 (Boston, MA, USA) (SIGIR ’09) . Association for Computing Machinery, New York, NY, USA, 758–759
Gordon V. Cormack, Charles L A Clarke, and Stefan Buettcher. 2009 · 2019
Earlier work this paper cites.
On Hallucination and Predictive Uncertainty in Conditional Language Generation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume . Association for Computational Linguistics, Online, 2734–2744
Yijun Xiao and William Yang Wang. 2021 · 2021
Earlier work this paper cites.
BARTScore: Evaluating Generated Text as Text Generation. In Advances in Neural Information Processing Systems , M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan (Eds.), Vol. 34. Curran Associates, Inc., 27263–27277
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Earlier work this paper cites.
Promptagator: Few-shot Dense Retrieval From 8 Examples
Zhuyun Dai, Vincent Y. Zhao, Ji Ma, Yi Luan, Jianmo Ni, Jing Lu, Anton Bakalov, Kelvin Guu, Keith B. Hall, and Ming-Wei Chang. 2022 · 2022
Earlier work this paper cites.
Efficient inference for Kendall’s tau
Samuel Perreault. 2022 · 2022
Earlier work this paper cites.
Knowledge of Knowledge: Exploring Known-Unknowns Uncertainty with Large Language Models
Alfonso Amayuelas, Liangming Pan, Wenhu Chen, and William Wang. 2023 · 2023
Cited alongside, same era.
Anastasios N. Angelopoulos, Stephen Bates, Clara Fannjiang, Michael I. Jordan, and Tijana Zrnic. 2023 · 2023
Cited alongside, same era.
The Internal State of an LLM Knows When It’s Lying
Amos Azaria and Tom Mitchell. 2023 · 2023
Cited alongside, same era.
Benchmarking Large Language Models in Retrieval-Augmented Generation
Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun. 2023 · 2023
Cited alongside, same era.
scipy.stats.kendalltau
[n. d.] · 2024
Closest in time.
The Claude 3 Model Family: Opus, Sonnet, Haiku
Anthropic. 2024 · 2024
Closest in time.
A Comparison of Methods for Evaluating Generative IR
Negar Arabzadeh and Charles L. A. Clarke. 2024 · 2024
Closest in time.
ARAGOG: Advanced RAG Output Grading
Matouš Eibich, Shivay Nagpal, and Alexander Fred-Ojala. 2024 · 2024
Closest in time.
Toward Optimising a Retrieval Augmented Generation Pipeline using Large Language Model
Gentrit Fazlija. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nicholas D. Edwards, Enzo de Jong, and Stephen T. Ferguson. 2023 · 2023
Cited alongside, same era.
RAGAS: Automated Evaluation of Retrieval Augmented Generation
Shahul Es, Jithin James, Luis Espinosa-Anke, and Steven Schockaert. 2023 · 2023
Cited alongside, same era.
Perspectives on Large Language Models for Relevance Judgment. In Proceedings of the 2023 ACM SIGIR International Conference on Theory of Information Retrieval (ICTIR ’23) . ACM
Guglielmo Faggioli, Laura Dietz, Charles L. A. Clarke, Gianluca Demartini, Matthias Hagen, Claudia Hauff, Noriko Kando, Evangelos Kanoulas, Martin Potthast, Benno Stein, and Henning Wachsmuth. 2023 · 2023
Cited alongside, same era.
InPars-v2: Large Language Models as Efficient Dataset Generators for Information Retrieval
Vitor Jeronymo, Luiz Bonifacio, Hugo Abonizio, Marzieh Fadaee, Roberto Lotufo, Jakub Zavrel, and Rodrigo Nogueira. 2023 · 2023
Cited alongside, same era.
Survey of Hallucination in Natural Language Generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023 · 2023
Cited alongside, same era.
SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
Potsawee Manakul, Adian Liusie, and Mark J. F. Gales. 2023 · 2023
Cited alongside, same era.
Large language models can accurately predict searcher preferences
Paul Thomas, Seth Spielman, Nick Craswell, and Bhaskar Mitra. 2023 · 2023
Cited alongside, same era.
Do Large Language Models Know What They Don’t Know?. In Findings of the Association for Computational Linguistics: ACL 2023 . 8653–8665
Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Xuan-Jing Huang. 2023 · 2023
Cited alongside, same era.
Hui Huang, Yingqi Qu, Jing Liu, Muyun Yang, and Tiejun Zhao. 2024 · 2024
Closest in time.
A Comprehensive Survey of Evaluation Techniques for Recommendation Systems
Aryan Jadon and Avinash Patil. 2024 · 2024
Closest in time.
FaaF: Facts as a Function for the evaluation of generated text
Vasileios Katranidis and Gabor Barany. 2024 · 2024
Closest in time.
GPT-4 Turbo and GPT-4
OpenAI. 2024 · 2024
Closest in time.
Rag-Fusion: A New Take on Retrieval Augmented Generation
Zackary Rackauckas. 2024 · 2024
Closest in time.
ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems
Jon Saad-Falcon, Omar Khattab, Christopher Potts, and Matei Zaharia. 2024 · 2024
Closest in time.
Multilingual E5 Text Embeddings: A Technical Report
Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. 2024 · 2024
Closest in time.
Corrective Retrieval Augmented Generation
Shi-Qi Yan, Jia-Chen Gu, Yun Zhu, and Zhen-Hua Ling. 2024 · 2024
Closest in time.