Fetching the paper…
Reading the bibliography…
We propose a new method to measure the task-specific accuracy of Retrieval-Augmented Large Language Models (RAG).
Generalization through memorization: Nearest neighbor language models
Khandelwal, U., Levy, O., Jurafsky, D., Zettlemoyer, L., and Lewis, M · 1911
Earlier work this paper cites.
Taxonomy of educational objectives: The classification of educational goals. Handbook 1: Cognitive domain
Bloom, B. S., Engelhart, M. D., Furst, E. J., Hill, W. H., and Krathwohl, D. R · 1956
Earlier work this paper cites.
Studies in mathematical psychology: I. probabilistic models for some intelligence and attainment tests
Rasch, G · 1960
Earlier work this paper cites.
Statistical theories of mental test scores
Lord, F., Novick, M., and Birnbaum, A · 1968
Earlier work this paper cites.
Fundamentals of item response theory , volume 2
Hambleton, R. K., Swaminathan, H., and Rogers, H. J · 1991
Earlier work this paper cites.
REALM: retrieval-augmented language model pre-training
Guu, K., Lee, K., Tung, Z., Pasupat, P., and Chang, M · 2002
Earlier work this paper cites.
A revision of bloom’s taxonomy: An overview
Krathwohl, D. R · 2002
Earlier work this paper cites.
Machine translation evaluation: N-grams to the rescue
Papineni, K · 2002
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive NLP tasks
Lewis, P. S. H., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., and Kiela, D · 2005
Earlier work this paper cites.
Natural language processing with Python: analyzing text with the natural language toolkit
Bird, S., Klein, E., and Loper, E · 2009
Earlier work this paper cites.
The probabilistic relevance framework: Bm25 and beyond
Robertson, S., Zaragoza, H., et al · 2009
Earlier work this paper cites.
Item response theory
Embretson, S. E. and Reise, S. P · 2013
Earlier work this paper cites.
Siamese neural networks for one-shot image recognition
Koch, G., Zemel, R., Salakhutdinov, R., et al · 2015
Earlier work this paper cites.
Item response theory
Cai, L., Choi, K., Hansen, M., and Harrell, L · 2016
Earlier work this paper cites.
Making sense of item response theory in machine learning
Martínez-Plumed, F., Prudêncio, R. B., Martínez-Usó, A., and Hernández-Orallo, J · 2016
Earlier work this paper cites.
Why we need new evaluation metrics for nlg
Novikova, J., Dušek, O., Curry, A. C., and Rieser, V · 2017
Earlier work this paper cites.
Item response theory in ai: Analysing machine learning classifiers at the instance level
Martínez-Plumed, F., Prudêncio, R. B., Martínez-Usó, A., and Hernández-Orallo, J · 2019
Cited alongside, same era.
Multiqa: An empirical investigation of generalization and transfer in reading comprehension
Talmor, A. and Berant, J · 2019
Cited alongside, same era.
Deep-irt: Make deep learning based knowledge tracing explainable using item response theory
Yeung, C.-K · 2019
Cited alongside, same era.
Climbing towards nlu: On meaning, form, and understanding in the age of data
Bender, E. M. and Koller, A · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Evaluating correctness and faithfulness of instruction-following models for question answering
Adlakha, V., BehnamGhader, P., Lu, X. H., Meade, N., and Reddy, S · 2023
Later among the works it cites.
Benchmarking large language models in retrieval-augmented generation
Chen, J., Lin, H., Han, X., and Sun, L · 2023
Later among the works it cites.
Ragas: Automated evaluation of retrieval augmented generation
Es, S., James, J., Espinosa-Anke, L., and Schockaert, S · 2023
Later among the works it cites.
Gptscore: Evaluate as you desire
Fu, J., Ng, S.-K., Jiang, Z., and Liu, P · 2023
Later among the works it cites.
Ralle: A framework for developing and evaluating retrieval-augmented large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Twenty years of confusion in human evaluation: Nlg needs evaluation sheets and standardised definitions
Howcroft, D. M., Belz, A., Clinciu, M., Gkatzia, D., Hasan, S. A., Mahamood, S., Mille, S., Van Miltenburg, E., Santhanam, S., and Rieser, V · 2020
Cited alongside, same era.
Item response theory for efficient human evaluation of chatbots
Sedoc, J. and Ungar, L · 2020
Cited alongside, same era.
Minilm: deep self-attention distillation for task-agnostic compression of pre-trained transformers
Wang, W., Wei, F., Dong, L., Bao, H., Yang, N., and Zhou, M · 2020
Cited alongside, same era.
Improving language models by retrieving from trillions of tokens
Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Millican, K., van den Driessche, G., Lespiau, J., Damoc, B., Clark, A., de Las Casas, D., Guy, A., Menick, J., Ring, R., Hennigan, T., Huang, S., Maggiore, L., Jones, C., Cassirer, A., Brock, A., Paganini, M., Irving, G., Vinyals, O., Osindero, S., Simonyan, K., Rae, J. W., Elsen, E., and Sifre, L · 2021
Cited alongside, same era.
What will it take to fix benchmarking in natural language understanding?
Bowman, S. R. and Dahl, G. E · 2021
Cited alongside, same era.
A statistical analysis of summarization evaluation metrics using resampling methods
Deutsch, D., Dror, R., and Roth, D · 2021
Cited alongside, same era.
Summeval: Re-evaluating summarization evaluation
Fabbri, A. R., Kryściński, W., McCann, B., Xiong, C., Socher, R., and Radev, D · 2021
Cited alongside, same era.
Hoshi, Y., Miyashita, D., Ng, Y., Tatsuno, K., Morioka, Y., Torii, O., and Deguchi, J · 2023
Later among the works it cites.
Mistral 7b, 2023
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Later among the works it cites.
Hagrid: A human-llm collaborative dataset for generative information-seeking with attribution
Kamalloo, E., Jafari, A., Zhang, X., Thakur, N., and Lin, J · 2023
Later among the works it cites.
What we evaluate when we evaluate recommender systems: Understanding recommender systems’ performance using item response theory
Liu, Y., Medlar, A., and Glowacka, D · 2023
Later among the works it cites.
Factscore: Fine-grained atomic evaluation of factual precision in long form text generation
Min, S., Krishna, K., Lyu, X., Lewis, M., Yih, W.-t., Koh, P. W., Iyyer, M., Zettlemoyer, L., and Hajishirzi, H · 2023
Later among the works it cites.
Generating benchmarks for factuality evaluation of language models
Muhlgay, D., Ram, O., Magar, I., Levine, Y., Ratner, N., Belinkov, Y., Abend, O., Leyton-Brown, K., Shashua, A., and Shoham, Y · 2023
Later among the works it cites.
Ares: An automated evaluation framework for retrieval-augmented generation systems
Saad-Falcon, J., Khattab, O., Potts, C., and Zaharia, M · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T · 2023
Later among the works it cites.
Interactive visual reasoning under uncertainty
Xu, M., Jiang, G., Liang, W., Zhang, C., and Zhu, Y · 2023
Later among the works it cites.
Kola: Carefully benchmarking world knowledge of large language models
Yu, J., Wang, X., Tu, S., Cao, S., Zhang-Li, D., Lv, X., Peng, H., Yao, Z., Zhang, X., Li, H., et al · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2023
Later among the works it cites.