Fetching the paper…
Reading the bibliography…
The conventional paradigm of using large language models (LLMs) for natural language generation (NLG) evaluation relies on pre-defined task definitions and evaluation criteria, positioning LLMs as "passive critics" that strictly follow developer-provided guidelines.
A new measure of rank correlation
Maurice G Kendall. 1938 · 1938
Earlier work this paper cites.
Evaluation in the context of natural language generation
Chris Mellish and Robert Dale. 1998 · 1998
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Spearman rank correlation
Jerrold H Zar. 2005 · 2005
Earlier work this paper cites.
Evaluation of text generation: A survey
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020 · 2006
Earlier work this paper cites.
GLEU: Automatic evaluation of sentence-level fluency
Andrew Mutton, Mark Dras, Stephen Wan, and Robert Dale. 2007 · 2007
Earlier work this paper cites.
The meteor metric for automatic evaluation of machine translation
Alon Lavie and Michael J Denkowski. 2009 · 2009
Earlier work this paper cites.
Computing krippendorff’s alpha-reliability
Klaus Krippendorff. 2011 · 2011
Earlier work this paper cites.
A guide to appropriate use of correlation coefficient in medical research
Mavuto M Mukaka. 2012 · 2012
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2015 · 2015
Earlier work this paper cites.
Semantically conditioned lstm-based natural language generation for spoken dialogue systems
Tsung-Hsien Wen, Milica Gasic, Nikola Mrkšić, Pei-Hao Su, David Vandyke, and Steve Young. 2015 · 2015
Earlier work this paper cites.
Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies
Max Grusky, Mor Naaman, and Yoav Artzi. 2018 · 2018
Earlier work this paper cites.
Sentence-level fluency evaluation: References help, but can be spared!
Katharina Kann, Sascha Rothe, and Katja Filippova. 2018 · 2018
Earlier work this paper cites.
TIGEr: Text-to-image grounding for image caption evaluation
Ming Jiang, Qiuyuan Huang, Lei Zhang, Xin Wang, Pengchuan Zhang, Zhe Gan, Jana Diesner, and Jianfeng Gao. 2019 · 2019
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 2019
Earlier work this paper cites.
Usr: An unsupervised and reference free evaluation metric for dialog generation
Shikib Mehri and Maxine Eskenazi. 2020 · 2020
Earlier work this paper cites.
All that’s ‘human’is not gold: Evaluating human evaluation of generated text
Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A Smith. 2021 · 2021
Cited alongside, same era.
Summeval: Re-evaluating summarization evaluation
Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021 · 2021
Cited alongside, same era.
Openmeva: A benchmark for evaluating open-ended story generation metrics
Jian Guan, Zhexin Zhang, Zhuoer Feng, Zitao Liu, Wenbiao Ding, Xiaoxi Mao, Changjie Fan, and Minlie Huang. 2021 · 2021
Cited alongside, same era.
Bartscore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022 · 2022
Cited alongside, same era.
G-eval: Nlg evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023a · 2023
Later among the works it cites.
Towards interpretable and efficient automatic reference-based summarization evaluation
Yixin Liu, Alexander Richard Fabbri, Yilun Zhao, Pengfei Liu, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, and Dragomir Radev. 2023b · 2023
Later among the works it cites.
Exploring prompting large language models as explainable metrics
Ghazaleh Mahmoudi. 2023 · 2023
Later among the works it cites.
Automatic prompt optimization with" gradient descent" and beam search
Reid Pryzant, Dan Iter, Jerry Li, Yin Tat Lee, Chenguang Zhu, and Michael Zeng. 2023 · 2023
Later among the works it cites.
Evaluating evaluation metrics: A framework for analyzing nlg evaluation metrics using measurement theory
Ziang Xiao, Susu Zhang, Vivian Lai, and Q Vera Liao. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han. 2022 · 2022
Cited alongside, same era.
A closer look into using large language models for automatic evaluation
Cheng-Han Chiang and Hung-yi Lee. 2023b · 2023
Cited alongside, same era.
Coascore: Chain-of-aspects prompting for nlg evaluation
Peiyuan Gong and Jiaxin Mao. 2023 · 2023
Cited alongside, same era.
Zero-shot faithfulness evaluation for text summarization with foundation language model
Qi Jia, Siyu Ren, Yizhu Liu, and Kenny Zhu. 2023 · 2023
Cited alongside, same era.
Tigerscore: Towards building explainable metric for all text generation tasks
Dongfu Jiang, Yishan Li, Ge Zhang, Wenhao Huang, Bill Yuchen Lin, and Wenhu Chen. 2023 · 2023
Cited alongside, same era.
Pei Ke, Bosi Wen, Zhuoer Feng, Xiao Liu, Xuanyu Lei, Jiale Cheng, Shengyuan Wang, Aohan Zeng, Yuxiao Dong, Hongning Wang, et al. 2023 · 2023
Cited alongside, same era.
Dspy: Compiling declarative language model calls into self-improving pipelines
Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav Santhanam, Sri Vardhamanan, Saiful Haq, Ashutosh Sharma, Thomas T. Joshi, Hanna Moazam, Heather Miller, Matei Zaharia, and Christopher Potts. 2023 · 2023
Cited alongside, same era.
Instructscore: Towards explainable text generation evaluation with automatic feedback
Wenda Xu, Danqing Wang, Liangming Pan, Zhenqiao Song, Markus Freitag, William Wang, and Lei Li. 2023 · 2023
Later among the works it cites.
Large language models as optimizers
Chengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu, Quoc V Le, Denny Zhou, and Xinyun Chen. 2023 · 2023
Later among the works it cites.
Batcheval: Towards human-like text evaluation
Peiwen Yuan, Shaoxiong Feng, Yiwei Li, Xinglin Wang, Boyuan Pan, Heda Wang, and Kan Li. 2023 · 2023
Later among the works it cites.
Large language models are human-level prompt engineers
Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. 2023 · 2023
Later among the works it cites.
Evaluating correctness and faithfulness of instruction-following models for question answering
Vaibhav Adlakha, Parishad BehnamGhader, Xing Han Lu, Nicholas Meade, and Siva Reddy. 2024 · 2024
Closest in time.
Gptscore: Evaluate as you desire
Jinlan Fu, See Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2024 · 2024
Closest in time.
Llm-based nlg evaluation: Current status and challenges
Mingqi Gao, Xinyu Hu, Jie Ruan, Xiao Pu, and Xiaojun Wan. 2024 · 2024
Closest in time.
Themis: A reference-free nlg evaluation language model with flexibility and interpretability
Xinyu Hu, Li Lin, Mingqi Gao, Xunjian Yin, and Xiaojun Wan. 2024 · 2024
Closest in time.
Intent-based prompt calibration: Enhancing prompt optimization with synthetic boundary cases
Elad Levi, Eli Brosh, and Matan Friedmann. 2024 · 2024
Closest in time.
Calibrating llm-based evaluator
Yuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, and Qi Zhang. 2024a · 2024
Closest in time.
Optimizing instructions and demonstrations for multi-stage language model programs
Krista Opsahl-Ong, Michael J Ryan, Josh Purtell, David Broman, Christopher Potts, Matei Zaharia, and Omar Khattab. 2024 · 2024
Closest in time.
Assessing the human likeness of AI-generated counterspeech
Xiaoying Song, Sujana Mamidisetty, Eduardo Blanco, and Lingzi Hong. 2025 · 2025
Closest in time.