Fetching the paper…
Reading the bibliography…
Evaluating the quality of text generated by large language models (LLMs) remains a significant challenge.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. (2002) · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y. (2004) · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Banerjee, S. and Lavie, A. (2005) · 2005
Earlier work this paper cites.
SummEval: Re-evaluating Summarization Evaluation
Fabbri, A. R., Kryściński, W., McCann, B., Xiong, C., Socher, R., and Radev, D. (2021) · 2021
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., ichter, b., Xia, F., Chi, E., Le, Q. V., and Zhou, D. (2022) · 2022
Cited alongside, same era.
Gptscore: Evaluate as you desire
Fu, J., Ng, S.-K., Jiang, Z., and Liu, P. (2023) · 2023
Cited alongside, same era.
G-eval: Nlg evaluation using gpt-4 with better human alignment
Liu, Y., Iter, D., Xu, Y., Wang, S., Xu, R., and Zhu, C. (2023) · 2023
Later among the works it cites.
Datasets for portuguese legal semantic textual similarity
da Silva Junior, D., dos Santos Corval, P. R., de Oliveira, D., and Paes, A. (2024) · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…