Fetching the paper…
Reading the bibliography…
Despite Retrieval-Augmented Generation (RAG) showing promising capability in leveraging external knowledge, a comprehensive evaluation of RAG systems is still challenging due to the modular nature of RAG, evaluation of long-form responses and reliability of measurements.
The trec-8 question answering track report
E. M. Voorhees et al · 1999
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
The probabilistic relevance framework: Bm25 and beyond
S. Robertson, H. Zaragoza, et al · 2009
Earlier work this paper cites.
An overview of the bioasq large-scale biomedical semantic indexing and question answering competition
G. Tsatsaronis, G. Balikas, P. Malakasiotis, I. Partalas, M. Zschunke, M. R. Alvers, D. Weissenborn, A. Krithara, S. Petridis, D. Polychronopoulos, et al · 2015
Earlier work this paper cites.
Www’18 open challenge: financial opinion mining and question answering
M. Maia, S. Handschuh, A. Freitas, B. Davis, R. McDermott, M. Zarrouk, and A. Balahur · 2018
Earlier work this paper cites.
Natural questions: A benchmark for question answering research
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, K. Toutanova, L. Jones, M. Kelcey, M.-W. Chang, A. M. Dai, J. Uszkoreit, Q. Le, and S. Petrov · 2019
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi · 2019
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, et al · 2020
Earlier work this paper cites.
S2ORC: The semantic scholar open research corpus
K. Lo, L. L. Wang, M. Neumann, R. Kinney, and D. Weld · 2020
Earlier work this paper cites.
Generation-augmented retrieval for open-domain question answering
Y. Mao, P. He, X. Liu, Y. Shen, J. Gao, J. Han, and W. Chen · 2020
Earlier work this paper cites.
Internet-augmented dialogue generation
M. Komeili, K. Shuster, and J. Weston · 2021
Earlier work this paper cites.
Retrieval augmented code generation and summarization
M. R. Parvez, W. U. Ahmad, S. Chakraborty, B. Ray, and K.-W. Chang · 2021
Earlier work this paper cites.
Retrieval augmentation reduces hallucination in conversation
K. Shuster, S. Poff, M. Chen, D. Kiela, and J. Weston · 2021
Earlier work this paper cites.
Efficient retrieval augmented generation from unstructured knowledge for task-oriented dialog
D. Thulke, N. Daheim, C. Dugast, and H. Ney · 2021
Earlier work this paper cites.
LangChain
H. Chase · 2022
Earlier work this paper cites.
Text and code embeddings by contrastive pre-training
A. Neelakantan, T. Xu, R. Puri, A. Radford, J. M. Han, J. Tworek, Q. Yuan, N. Tezak, J. W. Kim, C. Hallacy, et al · 2022
Earlier work this paper cites.
ColBERTv2: Effective and efficient retrieval via lightweight late interaction
K. Santhanam, O. Khattab, J. Saad-Falcon, C. Potts, and M. Zaharia · 2022
Cited alongside, same era.
Docprompting: Generating code by retrieving the docs
S. Zhou, U. Alon, F. F. Xu, Z. Wang, Z. Jiang, and G. Neubig · 2022
Cited alongside, same era.
Retrieval-based language models and applications
A. Asai, S. Min, Z. Zhong, and D. Chen · 2023
Cited alongside, same era.
Ragas: Automated evaluation of retrieval augmented generation
S. Es, J. James, L. Espinosa-Anke, and S. Schockaert · 2023
Cited alongside, same era.
Retrieval-augmented generation for large language models: A survey
Y. Gao, Y. Xiong, X. Gao, K. Jia, J. Pan, Y. Bi, Y. Dai, J. Sun, and H. Wang · 2023
Cited alongside, same era.
Refchecker: Reference-based fine-grained hallucination checker and benchmark for large language models
X. Hu, D. Ru, L. Qiu, Q. Guo, T. Zhang, Y. Xu, Y. Luo, P. Liu, Y. Zhang, and Z. Zhang · 2024
Closest in time.
A survey on retrieval-augmented text generation for large language models
Y. Huang and J. Huang · 2024
Closest in time.
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. d. l. Casas, E. B. Hanna, F. Bressand, et al · 2024
Closest in time.
Faaf: Facts as a function for the evaluation of rag systems
V. Katranidis and G. Barany · 2024
Closest in time.
Interpretable long-form legal question answering with retrieval-augmented large language models
A. Louis, G. van Dijck, and G. Spanakis · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Retrieval-augmented code generation for universal information extraction
Y. Guo, Z. Li, X. Jin, Y. Liu, Y. Zeng, W. Liu, X. Li, P. Yang, L. Bai, J. Guo, et al · 2023
Cited alongside, same era.
Robustqa: Benchmarking the robustness of domain adaptation for open-domain question answering
R. Han, P. Qi, Y. Zhang, L. Liu, J. Burger, W. Wang, Z. Huang, B. Xiang, and D. Roth · 2023
Cited alongside, same era.
L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin, et al · 2023
Cited alongside, same era.
Evaluating verifiability in generative search engines
N. F. Liu, T. Zhang, and P. Liang · 2023
Cited alongside, same era.
Recall: A benchmark for llms robustness against external counterfactual knowledge
Y. Liu, L. Huang, S. Li, S. Chen, H. Zhou, F. Meng, J. Zhou, and X. Sun · 2023
Cited alongside, same era.
Ares: An automated evaluation framework for retrieval-augmented generation systems, 2023
J. Saad-Falcon, O. Khattab, C. Potts, and M. Zaharia · 2023
Cited alongside, same era.
Nomiracl: Knowing when you don’t know for robust multilingual retrieval-augmented generation
N. Thakur, L. Bonifacio, X. Zhang, O. Ogundepo, E. Kamalloo, D. Alfonso-Hermelo, X. Li, Q. Liu, B. Chen, M. Rezagholizadeh, et al · 2023
Cited alongside, same era.
Closest in time.
Y. Lyu, Z. Li, S. Niu, F. Xiong, B. Tang, W. Wang, H. Wu, H. Liu, T. Xu, and E. Chen · 2024
Closest in time.
Clapnq: Cohesive long-form answers from passages in natural questions for rag systems, 2024
S. Rosenthal, A. Sil, R. Florian, and S. Roukos · 2024
Closest in time.
Prompt-based code completion via multi-retrieval augmented generation
H. Tan, Q. Luo, L. Jiang, Z. Zhan, J. Li, H. Zhang, and Y. Zhang · 2024
Closest in time.
Multihop-rag: Benchmarking retrieval-augmented generation for multi-hop queries
Y. Tang and Y. Yang · 2024
Closest in time.
A comprehensive survey of hallucination mitigation techniques in large language models
S. Tonmoy, S. Zaman, V. Jain, A. Rani, V. Rawte, A. Chadha, and A. Das · 2024
Closest in time.
Evaluating open-qa evaluation
C. Wang, S. Cheng, Q. Guo, Y. Yue, B. Ding, Z. Xu, Y. Wang, X. Hu, Z. Zhang, and Y. Zhang · 2024
Closest in time.
Novelqa: A benchmark for long-range novel question answering, 2024
C. Wang, R. Ning, B. Pan, T. Wu, Q. Guo, C. Deng, G. Bao, Q. Wang, and Y. Zhang · 2024
Closest in time.
How faithful are rag models? quantifying the tug-of-war between rag and llms’ internal prior
K. Wu, E. Wu, and J. Zou · 2024
Closest in time.
Benchmarking retrieval-augmented generation for medicine
G. Xiong, Q. Jin, Z. Lu, and A. Zhang · 2024
Closest in time.
Kiwi: A dataset of knowledge-intensive writing instructions for answering research questions, 2024
F. Xu, K. Lo, L. Soldaini, B. Kuehl, E. Choi, and D. Wadden · 2024
Closest in time.
Let llms take on the latest challenges! a chinese dynamic question answering benchmark
Z. Xu, Y. Li, R. Ding, X. Wang, B. Chen, Y. Jiang, X. Deng, J. Ma, H.-T. Zheng, W. Lu, et al · 2024
Closest in time.
Evaluation of retrieval-augmented generation: A survey
H. Yu, A. Gan, K. Zhang, S. Tong, Q. Liu, and Z. Liu · 2024
Closest in time.
Almanac—retrieval-augmented language models for clinical medicine
C. Zakka, R. Shad, A. Chaurasia, A. R. Dalal, J. L. Kim, M. Moor, R. Fong, C. Phillips, K. Alexander, E. Ashley, et al · 2024
Closest in time.