Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are increasingly serving as evaluators in Natural Language Generation (NLG) tasks; this is often referred to as ``LLM-as-a-judge'' paradigm.
Individual comparisons by ranking methods
F Wilcoxon. 1945 · 1945
Earlier work this paper cites.
Effects of culture and response format on extreme response style
C Harry Hui and Harry C Triandis. 1989 · 1989
Earlier work this paper cites.
Multidimensional quality metrics: a flexible system for assessing translation quality
Aljoscha Burchardt. 2013 · 2013
Earlier work this paper cites.
Response style behavior: question format dependent or personal style?
Natalia D Kieruj and Guy Moors. 2013 · 2013
Earlier work this paper cites.
Response styles in survey research: A literature review of antecedents, consequences, and remedies
Yves Van Vaerenbergh and Troy D Thomas. 2013 · 2013
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Lsdsem 2017 shared task: The story cloze test
Nasrin Mostafazadeh, Michael Roth, Annie Louis, Nathanael Chambers, and James Allen. 2017 · 2017
Earlier work this paper cites.
Rankme: Reliable human ratings for natural language generation
Jekaterina Novikova, Ondřej Dušek, and Verena Rieser. 2018 · 2018
Earlier work this paper cites.
The harmonic mean p-value for combining dependent tests
Daniel J Wilson. 2019 · 2019
Earlier work this paper cites.
“this is a problem, don’t you agree?” framing and bias in human evaluation for natural language generation
Stephanie Schoch, Diyi Yang, and Yangfeng Ji. 2020 · 2020
Earlier work this paper cites.
Combining p-values via averaging
Vladimir Vovk and Ruodu Wang. 2020 · 2020
Earlier work this paper cites.
Automatic chain of thought prompting in large language models
Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2022 · 2020
Earlier work this paper cites.
Summeval: Re-evaluating summarization evaluation
Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021 · 2021
Earlier work this paper cites.
GO FIGURE: A meta evaluation of factuality in summarization
Saadia Gabriel, Asli Celikyilmaz, Rahul Jha, Yejin Choi, and Jianfeng Gao. 2021 · 2021
Earlier work this paper cites.
Sumpubmed: Summarization dataset of pubmed scientific article
Vivek Gupta, Prerna Bharti, Pegah Nokhiz, and Harish Karnick. 2020 · 2021
Cited alongside, same era.
Perturbation checklists for evaluating nlg evaluation metrics
Ananya B Sai, Tanay Dixit, Dev Yashpal Sheth, Sreyas Mohan, and Mitesh M Khapra. 2021 · 2021
Cited alongside, same era.
Tomayto, tomahto. beyond token-level answer equivalence for question answering evaluation
Jannis Bulian, Christian Buck, Wojciech Gajewski, Benjamin Börschinger, and Tal Schuster. 2022 · 2022
Cited alongside, same era.
Findings of the 2022 conference on machine translation (wmt22)
Tom Kocmi, Rachel Bawden, Ondřej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Thamme Gowda, Yvette Graham, Roman Grundkiewicz, Barry Haddow, et al. 2022 · 2022
Cited alongside, same era.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. 2023 · 2023
Cited alongside, same era.
Is chatgpt a good nlg evaluator? a preliminary study
Jiaan Wang, Yunlong Liang, Fandong Meng, Zengkui Sun, Haoxiang Shi, Zhixu Li, Jinan Xu, Jianfeng Qu, and Jie Zhou. 2023 · 2023
Later among the works it cites.
Large language models for healthcare data augmentation: An example on patient-trial matching
Jiayi Yuan, Ruixiang Tang, Xiaoqian Jiang, and Xia Hu. 2023 · 2023
Later among the works it cites.
Wider and deeper llm networks are fairer llm evaluators
Xinghua Zhang, Bowen Yu, Haiyang Yu, Yangyu Lv, Tingwen Liu, Fei Huang, Hongbo Xu, and Yongbin Li. 2023 · 2023
Later among the works it cites.
Spec: a soft prompt-based calibration on performance variability of large language model in clinical notes summarization
Yu-Neng Chuang, Ruixiang Tang, Xiaoqian Jiang, and Xia Hu. 2024 · 2024
Closest in time.
Evalullm: Llm assisted evaluation of generative outputs
Michael Desmond, Zahra Ashktorab, Qian Pan, Casey Dugan, and James M Johnson. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chateval: Towards better llm-based evaluators through multi-agent debate
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2023 · 2023
Cited alongside, same era.
Can large language models be an alternative to human evaluations?
Cheng-Han Chiang and Hung-yi Lee. 2023 · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023 · 2023
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023 · 2023
Cited alongside, same era.
Knowledge distillation of llm for education
Ehsan Latif, Luyang Fang, Ping Ma, and Xiaoming Zhai. 2023 · 2023
Cited alongside, same era.
Prd: Peer rank and discussion improve large language model based evaluations
Ruosen Li, Teerth Patel, and Xinya Du. 2023 · 2023
Cited alongside, same era.
G-eval: Nlg evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023a · 2023
Cited alongside, same era.
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024 · 2024
Closest in time.
Llm-based nlg evaluation: Current status and challenges
Mingqi Gao, Xinyu Hu, Jie Ruan, Xiao Pu, and Xiaojun Wan. 2024 · 2024
Closest in time.
Are LLM-based evaluators confusing NLG quality criteria?
Xinyu Hu, Mingqi Gao, Sen Hu, Yang Zhang, Yicheng Chen, Teng Xu, and Xiaojun Wan. 2024 · 2024
Closest in time.
Leveraging large language models for nlg evaluation: A survey
Zhen Li, Xiaohan Xu, Tao Shen, Can Xu, Jia-Chen Gu, and Chongyang Tao. 2024 · 2024
Closest in time.
Llm comparative assessment: Zero-shot nlg evaluation through pairwise comparisons using large language models
Adian Liusie, Potsawee Manakul, and Mark Gales. 2024 · 2024
Closest in time.
Large language models show human-like social desirability biases in survey responses
Aadesh Salecha, Molly E Ireland, Shashanka Subrahmanya, João Sedoc, Lyle H Ungar, and Johannes C Eichstaedt. 2024 · 2024
Closest in time.
Judgebench: A benchmark for evaluating llm-based judges
Sijun Tan, Siyuan Zhuang, Kyle Montgomery, William Y Tang, Alejandro Cuadron, Chenguang Wang, Raluca Ada Popa, and Ion Stoica. 2024 · 2024
Closest in time.
Hui Wei, Shenghua He, Tian Xia, Andy Wong, Jingyang Lin, and Mei Han. 2024 · 2024
Closest in time.