Fetching the paper…
Reading the bibliography…
Since the adoption of large language models (LLMs) for text evaluation has become increasingly prevalent in the field of natural language processing (NLP), a series of existing works attempt to optimize the prompts for LLM evaluators to improve their alignment with human judgment.
The need for biases in learning generalizations
Tom M Mitchell. 1980 · 1980
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Tze Leung Lai and Herbert Robbins. 1985 · 1985
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Evaluation of text generation: A survey
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020 · 2006
Earlier work this paper cites.
Semantically conditioned LSTM-based natural language generation for spoken dialogue systems
Tsung-Hsien Wen, Milica Gašić, Nikola Mrkšić, Pei-Hao Su, David Vandyke, and Steve Young. 2015 · 2015
Earlier work this paper cites.
Topical-Chat: Towards Knowledge-Grounded Open-Domain Conversations
Karthik Gopalakrishnan, Behnam Hedayatnia, Qinlang Chen, Anna Gottardi, Sanjeev Kwatra, Anu Venkatesh, Raefer Gabriel, and Dilek Hakkani-Tür. 2019 · 2019
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Earlier work this paper cites.
SummEval: Re-evaluating summarization evaluation
Alexander R. Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021 · 2021
Earlier work this paper cites.
Of human criteria and automatic metrics: A benchmark of the evaluation of story generation
Cyril Chhun, Pierre Colombo, Fabian M. Suchanek, and Chloé Clavel. 2022 · 2022
Earlier work this paper cites.
Defining and characterizing reward gaming
Joar Skalse, Nikolaus Howe, Dmitrii Krasheninnikov, and David Krueger. 2022 · 2022
Earlier work this paper cites.
GPS: Genetic prompt search for efficient few-shot learning
Hanwei Xu, Yujun Chen, Yulun Du, Nan Shao, Wang Yanggang, Haiyu Li, and Zhilin Yang. 2022 · 2022
Earlier work this paper cites.
Exploring the use of large language models for reference-free text quality evaluation: An empirical study
Yi Chen, Rui Wang, Haiyun Jiang, Shuming Shi, and Ruifeng Xu. 2023 · 2023
Earlier work this paper cites.
A closer look into using large language models for automatic evaluation
Cheng-Han Chiang and Hung-yi Lee. 2023 · 2023
Earlier work this paper cites.
Multi-dimensional evaluation of text summarization with in-context learning
Sameer Jain, Vaishakh Keshava, Swarnashree Mysore Sathyendra, Patrick Fernandes, Pengfei Liu, Graham Neubig, and Chunting Zhou. 2023 · 2023
Earlier work this paper cites.
DecompEval: Evaluating generated texts as unsupervised decomposed question answering
Pei Ke, Fei Huang, Fei Mi, Yasheng Wang, Qun Liu, Xiaoyan Zhu, and Minlie Huang. 2023 · 2023
Earlier work this paper cites.
Which is better? exploring prompting strategy for LLM-based metrics
JoongHoon Kim, Sangmin Lee, Seung Hun Han, Saeran Park, Jiyoon Lee, Kiyoon Jeong, and Pilsung Kang. 2023 · 2023
Earlier work this paper cites.
Little giants: Exploring the potential of small LLMs as evaluation metrics in summarization in the Eval4NLP 2023 shared task
Neema Kotonya, Saran Krishnasamy, Joel Tetreault, and Alejandro Jaimes. 2023 · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Cited alongside, same era.
The language of prompting: What linguistic properties make a prompt successful?
Alina Leidinger, Robert van Rooij, and Ekaterina Shutova. 2023 · 2023
Cited alongside, same era.
Yen-Ting Lin and Yun-Nung Chen. 2023 · 2023
Cited alongside, same era.
G-eval: NLG evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Cited alongside, same era.
OpenAI. 2023 · 2023
Cited alongside, same era.
SocREval: Large language models with the socratic method for reference-free reasoning evaluation
Hangfeng He, Hongming Zhang, and Dan Roth. 2024 · 2024
Later among the works it cites.
Automatic engineering of long prompts
Cho-Jui Hsieh, Si Si, Felix Yu, and Inderjit Dhillon. 2024 · 2024
Later among the works it cites.
Themis: A reference-free NLG evaluation language model with flexibility and interpretability
Xinyu Hu, Li Lin, Mingqi Gao, Xunjian Yin, and Xiaojun Wan. 2024b · 2024
Later among the works it cites.
Hui Huang, Yingqi Qu, Jing Liu, Muyun Yang, and Tiejun Zhao. 2024 · 2024
Later among the works it cites.
CritiqueLLM: Towards an informative critique generation model for evaluation of large language model generation
Pei Ke, Bosi Wen, Andrew Feng, Xiao Liu, Xuanyu Lei, Jiale Cheng, Shengyuan Wang, Aohan Zeng, Yuxiao Dong, Hongning Wang, Jie Tang, and Minlie Huang. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
GrIPS: Gradient-free, edit-based instruction search for prompting large language models
Archiki Prasad, Peter Hase, Xiang Zhou, and Mohit Bansal. 2023 · 2023
Cited alongside, same era.
Automatic prompt optimization with “gradient descent” and beam search
Reid Pryzant, Dan Iter, Jerry Li, Yin Lee, Chenguang Zhu, and Michael Zeng. 2023 · 2023
Cited alongside, same era.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023 · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023 · 2023
Cited alongside, same era.
Survival of the most influential prompts: Efficient black-box prompt search via clustering and pruning
Han Zhou, Xingchen Wan, Ivan Vulić, and Anna Korhonen. 2023a · 2023
Cited alongside, same era.
Chateval: Towards better LLM-based evaluators through multi-agent debate
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2024 · 2024
Cited alongside, same era.
A survey on evaluation of large language models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. 2024 · 2024
Cited alongside, same era.
Later among the works it cites.
Generative judge for evaluating alignment
Junlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan, hai zhao, and Pengfei Liu. 2024 · 2024
Later among the works it cites.
Calibrating LLM-based evaluator
Yuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, and Qi Zhang. 2024c · 2024
Later among the works it cites.
Evaluating the evaluator: Measuring llms’ adherence to task evaluation instructions
Bhuvanashree Murugadoss, Christian Poelitz, Ian Drosos, Vu Le, Nick McKenna, Carina Suzana Negreanu, Chris Parnin, and Advait Sarkar. 2024 · 2024
Later among the works it cites.
Check-eval: A checklist-based approach for evaluating text quality
Jayr Pereira and Roberto Lotufo. 2024 · 2024
Later among the works it cites.
Is LLM-as-a-judge robust? investigating universal adversarial attacks on zero-shot LLM assessment
Vyas Raina, Adian Liusie, and Mark Gales. 2024 · 2024
Later among the works it cites.
One prompt to rule them all: LLMs for opinion summary evaluation
Tejpalsingh Siledar, Swaroop Nath, Sankara Muddu, Rupasai Rangaraju, Swaprava Nath, Pushpak Bhattacharyya, Suman Banerjee, Amey Patil, Sudhanshu Singh, Muthusamy Chelliah, and Nikesh Garera. 2024 · 2024
Later among the works it cites.
Large language models are inconsistent and biased evaluators
Rickard Stureborg, Dimitris Alikaniotis, and Yoshi Suhara. 2024 · 2024
Later among the works it cites.
Evaluating large language models at evaluating instruction following
Zhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng, Tanya Goyal, and Danqi Chen. 2024 · 2024
Later among the works it cites.
RewardBench: Evaluating reward models for language modeling
Nathan Lambert, Valentina Pyatkin, Jacob Morrison, LJ Miranda, Bill Yuchen Lin, Khyathi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, Noah A. Smith, and Hannaneh Hajishirzi. 2025 · 2025
Closest in time.
Can many-shot in-context learning help LLMs as evaluators? a preliminary empirical study
Mingyang Song, Mao Zheng, and Xuan Luo. 2025 · 2025
Closest in time.
Reviseval: Improving LLM-as-a-judge via response-adapted references
Qiyuan Zhang, Yufei Wang, Tiezheng YU, Yuxin Jiang, Chuhan Wu, Liangyou Li, Yasheng Wang, Xin Jiang, Lifeng Shang, Ruiming Tang, Fuyuan Lyu, and Chen Ma. 2025 · 2025
Closest in time.
Towards a unified multi-dimensional evaluator for text generation
Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han. 2022 · 2038
Closest in time.