Fetching the paper…
Reading the bibliography…
Retrieval-Augmented Generation (RAG) has become a standard architectural pattern for incorporating domain-specific knowledge into user-facing chat applications powered by Large Language Models (LLMs).
Ms marco: A human generated machine reading comprehension dataset
T. Nguyen, M. Rosenberg, X. Song, J. Gao, S. Tiwary, R. Majumder, and L. Deng · 2016
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and VERification
J. Thorne, A. Vlachos, C. Christodoulopoulos, and A. Mittal · 2018
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. W. Cohen, R. Salakhutdinov, and C. D. Manning · 2018
Earlier work this paper cites.
Wizard of wikipedia: Knowledge-powered conversational agents, 2019
E. Dinan, S. Roller, K. Shuster, A. Fan, M. Auli, and J. Weston · 2019
Earlier work this paper cites.
PubMedQA: A dataset for biomedical research question answering
Q. Jin, B. Dhingra, Z. Liu, W. Cohen, and X. Lu · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, M. Kelcey, J. Devlin, K. Lee, K. N. Toutanova, L. Jones, M.-W. Chang, A. Dai, J. Uszkoreit, Q. Le, and S. Petrov · 2019
Earlier work this paper cites.
Latent retrieval for weakly supervised open domain question answering
K. Lee, M.-W. Chang, and K. Toutanova · 2019
Earlier work this paper cites.
The TechQA dataset
V. Castelli, R. Chakravarti, S. Dana, A. Ferritto, R. Florian, M. Franz, D. Garg, D. Khandelwal, S. McCarley, M. McCawley, M. Nasr, L. Pan, C. Pendus, J. Pitrelli, S. Pujar, S. Roukos, A. Sakrajda, A. Sil, R. Uceda-Sosa, T. Ward, and R. Zhang · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, S. Riedel, and D. Kiela · 2020
Earlier work this paper cites.
COVID-QA: A question answering dataset for COVID-19
T. Möller, A. Reina, R. Jayakumar, and M. Pietsch · 2020
Earlier work this paper cites.
CORD-19: The COVID-19 open research dataset
L. L. Wang, K. Lo, Y. Chandrasekhar, R. Reas, J. Yang, D. Burdick, D. Eide, K. Funk, Y. Katsis, R. M. Kinney, Y. Li, Z. Liu, W. Merrill, P. Mooney, D. A. Murdick, D. Rishi, J. Sheehan, Z. Shen, B. Stilson, A. D. Wade, K. Wang, N. X. R. Wang, C. Wilhelm, B. Xie, D. M. Raymond, D. S. Weld, O. Etzioni, and S. Kohlmeier · 2020
Earlier work this paper cites.
FinQA: A dataset of numerical reasoning over financial data
Z. Chen, W. Chen, C. Smiley, S. Shah, I. Borova, D. Langdon, R. Moussa, M. Beane, T.-H. Huang, B. Routledge, and W. Y. Wang · 2021
Earlier work this paper cites.
Cuad: An expert-annotated nlp dataset for legal contract review
D. Hendrycks, C. Burns, A. Chen, and S. Ball · 2021
Earlier work this paper cites.
Question answering over electronic devices: A new benchmark dataset and a multi-task learning based QA framework
A. Nandy, S. Sharma, S. Maddhashiya, K. Sachdeva, P. Goyal, and N. Ganguly · 2021
Earlier work this paper cites.
KILT: a benchmark for knowledge intensive language tasks
F. Petroni, A. Piktus, A. Fan, P. Lewis, M. Yazdani, N. De Cao, J. Thorne, Y. Jernite, V. Karpukhin, J. Maillard, V. Plachouras, T. Rocktäschel, and S. Riedel · 2021
Earlier work this paper cites.
Increasing faithfulness in knowledge-grounded dialogue with controllable features
H. Rashkin, D. Reitter, G. S. Tomar, and D. Das · 2021
Earlier work this paper cites.
TAT-QA: A question answering benchmark on a hybrid of tabular and textual content in finance
F. Zhu, W. Lei, Y. Huang, C. Wang, S. Zhang, J. Lv, F. Feng, and T.-S. Chua · 2021
Cited alongside, same era.
Less annotating, more classifying – addressing the data scarcity issue of supervised machine learning with deep transfer learning and bert - nli
M. Laurer, W. van Atteveldt, A. Casas, and K. Welbers · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, b. ichter, F. Xia, E. Chi, Q. V. Le, and D. Zhou · 2022
Cited alongside, same era.
Making a miracl: Multilingual information retrieval across a continuum of languages, 2022
X. Zhang, N. Thakur, O. Ogundepo, E. Kamalloo, D. Alfonso-Hermelo, X. Li, Q. Liu, M. Rezagholizadeh, and J. Lin · 2022
Cited alongside, same era.
Evaluating correctness and faithfulness of instruction-following models for question answering
V. Adlakha, P. BehnamGhader, X. H. Lu, N. Meade, and S. Reddy · 2023
https://www.trulens.org/
Trulens, 2023 · 2023
Later among the works it cites.
Ragtruth: A hallucination corpus for developing trustworthy retrieval-augmented language models, 2023
Y. Wu, J. Zhu, S. Xu, K. Shum, C. Niu, R. Zhong, J. Song, and T. Zhang · 2023
Later among the works it cites.
Automatic evaluation of attribution by large language models
X. Yue, B. Wang, Z. Chen, K. Zhang, Y. Su, and H. Sun · 2023
Later among the works it cites.
Judging LLM-as-a-judge with MT-bench and chatbot arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica · 2023
Later among the works it cites.
RAGAs: Automated evaluation of retrieval augmented generation
S. Es, J. James, L. Espinosa Anke, and S. Schockaert · 2024
Closest in time.
A survey on retrieval-augmented text generation for large language models, 2024
Y. Huang and J. Huang · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Attributed question answering: Evaluation and modeling for attributed large language models, 2023
B. Bohnet, V. Q. Tran, P. Verga, R. Aharoni, D. Andor, L. B. Soares, M. Ciaramita, J. Eisenstein, K. Ganchev, J. Herzig, K. Hui, T. Kwiatkowski, J. Ma, J. Ni, L. S. Saralegui, T. Schuster, W. W. Cohen, M. Collins, D. Das, D. Metzler, S. Petrov, and K. Webster · 2023
Cited alongside, same era.
Benchmarking large language models in retrieval-augmented generation
J. Chen, H. Lin, X. Han, and L. Sun · 2023
Cited alongside, same era.
Felm: Benchmarking factuality evaluation of large language models
s. chen, Y. Zhao, J. Zhang, I.-C. Chern, S. Gao, P. Liu, and J. He · 2023
Cited alongside, same era.
The dangers of trusting stochastic parrots: Faithfulness and trust in open-domain conversational question answering
S. Chiesurin, D. Dimakopoulos, M. A. Sobrevilla Cabezudo, A. Eshghi, I. Papaioannou, V. Rieser, and I. Konstas · 2023
Cited alongside, same era.
Enabling large language models to generate text with citations
T. Gao, H. Yen, J. Yu, and D. Chen · 2023
Cited alongside, same era.
DeBERTav3: Improving deBERTa using ELECTRA-style pre-training with gradient-disentangled embedding sharing
P. He, J. Gao, and W. Chen · 2023
Cited alongside, same era.
Hagrid: A human-llm collaborative dataset for generative information-seeking with attribution, 2023
E. Kamalloo, A. Jafari, X. Zhang, N. Thakur, and J. Lin · 2023
Cited alongside, same era.
Closest in time.
Prometheus 2: An open source language model specialized in evaluating other language models, 2024
S. Kim, J. Suk, S. Longpre, B. Y. Lin, J. Shin, S. Welleck, G. Neubig, M. Lee, K. Lee, and M. Seo · 2024
Closest in time.
Attributionbench: How hard is automatic attribution evaluation?
Y. Li, X. Yue, Z. Liao, and H. Sun · 2024
Closest in time.
Chatqa: Building gpt-4 level conversational qa models
Z. Liu, W. Ping, R. Roy, P. Xu, C. Lee, M. Shoeybi, and B. Catanzaro · 2024
Closest in time.
Hallucination-free? assessing the reliability of leading ai legal research tools, 2024
V. Magesh, F. Surani, M. Dahl, M. Suzgun, C. D. Manning, and D. E. Ho · 2024
Closest in time.
Expertqa: Expert-curated questions and attributed answers, 2024
C. Malaviya, S. Lee, S. Chen, E. Sieber, M. Yatskar, and D. Roth · 2024
Closest in time.
Ares: An automated evaluation framework for retrieval-augmented generation systems
J. Saad-Falcon, O. Khattab, C. Potts, and M. Zaharia · 2024
Closest in time.
Domainrag: A chinese benchmark for evaluating domain-specific retrieval-augmented generation, 2024
S. Wang, J. Liu, S. Song, J. Cheng, Y. Fu, P. Guo, K. Fang, Y. Zhu, and Z. Dou · 2024
Closest in time.
Crag – comprehensive rag benchmark, 2024
X. Yang, K. Sun, H. Xin, Y. Sun, N. Bhalla, X. Chen, S. Choudhary, R. D. Gui, Z. W. Jiang, Z. Jiang, L. Kong, B. Moran, J. Wang, Y. E. Xu, A. Yan, C. Yang, E. Yuan, H. Zha, N. Tang, L. Chen, N. Scheffer, Y. Liu, N. Shah, R. Wanga, A. Kumar, W. tau Yih, and X. L. Dong · 2024
Closest in time.
FLASK: Fine-grained language model evaluation based on alignment skill sets
S. Ye, D. Kim, S. Kim, H. Hwang, S. Kim, Y. Jo, J. Thorne, J. Kim, and M. Seo · 2024
Closest in time.