Fetching the paper…
Reading the bibliography…
The effectiveness of automatic evaluation of generative models is typically measured by comparing the labels generated via automation with labels by humans using correlation metrics.
On information and sufficiency
Solomon Kullback and Richard A Leibler · 1951
Earlier work this paper cites.
A coefficient of agreement for nominal scales
Jacob Cohen · 1960
Earlier work this paper cites.
Measuring nominal scale agreement among many raters
Joseph L. Fleiss · 1971
Earlier work this paper cites.
Judgment under uncertainty: Heuristics and biases
Amos Tversky and Daniel Kahneman · 1974
Earlier work this paper cites.
Variants of uncertainty
Daniel Kahneman and Amos Tversky · 1982
Earlier work this paper cites.
Divergence measures based on the shannon entropy
J. Lin · 1991
Earlier work this paper cites.
Reliability in content analysis: Some common misconceptions and recommendations
Klaus Krippendorff · 2004
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Free-marginal multirater kappa (multirater k [free]): An alternative to fleiss’ fixed-marginal multirater kappa
Justus J Randolph · 2005
Earlier work this paper cites.
I like it… i like it not: Evaluating user ratings noise in recommender systems
Xavier Amatriain, Josep M. Pujol, and Nuria Oliver · 2009
Earlier work this paper cites.
Computing krippendorff’s alpha-reliability
Klaus Krippendorff · 2011
Earlier work this paper cites.
Interrater reliability: the kappa statistic
Mary L McHugh · 2012
Earlier work this paper cites.
Statistics corner: A guide to appropriate use of correlation coefficient in medical research
M M Mukaka · 2012
Earlier work this paper cites.
Assumptions behind intercoder reliability indices
Jun S. Liu Xinshu Zhao and Ke Deng · 2013
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning · 2015
Earlier work this paper cites.
Likert scale: Explored and explained
Ankur Joshi, Saket Kale, Satish Chandel, and Dinesh Pal · 2015
Earlier work this paper cites.
(re) visualizing rater agreement: Beyond single-parameter measures
David Eubanks · 2017
Cited alongside, same era.
Predicting aesthetic score distribution through cumulative jensen-shannon divergence
Xin Jin, Le Wu, Xiaodong Li, Siyu Chen, Siwei Peng, Jingying Chi, Shiming Ge, Chenggen Song, and Geng Zhao · 2018
Cited alongside, same era.
Scales of measurement and presentation of statistical data
Prabhaker Mishra, C M Pandey, Uttam Singh, and Anshul Gupta · 2018
Cited alongside, same era.
On the usefulness of interrater reliability coefficients
Debby ten Hove, Terrence D. Jorgensen, and L. Andries van der Ark · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Cited alongside, same era.
Agreement is overrated: A plea for correlation to assess human evaluation reliability
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Later among the works it cites.
Re-examining system-level correlations of automatic summarization evaluation metrics
Daniel Deutsch, Rotem Dror, and Dan Roth · 2022
Later among the works it cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al · 2022
Later among the works it cites.
DICES dataset: Diversity in conversational AI evaluation for safety
Lora Aroyo, Alex Taylor, Mark Diaz, Christopher M Homan, Alicia Parrish, Greg Serapio-Garcia, Vinodkumar Prabhakaran, and Ding Wang · 2023
Later among the works it cites.
A closer look into using large language models for automatic evaluation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jacopo Amidei, Paul Piwek, and Alistair Willis · 2019
Cited alongside, same era.
What Is the Best Response Scale for Survey and Questionnaire Design; Review of Different Lengths of Rating Scale / Attitude Scale / Likert Scale
Hamed Taherdoost · 2019
Cited alongside, same era.
Uncertain natural language inference
Tongfei Chen, Zhengping Jiang, Adam Poliak, Keisuke Sakaguchi, and Benjamin Van Durme · 2020
Cited alongside, same era.
USR: An unsupervised and reference free evaluation metric for dialog generation
Shikib Mehri and Maxine Eskenazi · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Cited alongside, same era.
Human evaluation of automatically generated text: Current trends and best practice guidelines
Chris van der Lee, Albert Gatt, Emiel van Miltenburg, and Emiel Krahmer · 2020
Cited alongside, same era.
Asking and answering questions to evaluate the factual consistency of summaries
Alex Wang, Kyunghyun Cho, and Mike Lewis · 2020
Cited alongside, same era.
Cheng-Han Chiang and Hung-yi Lee · 2023
Later among the works it cites.
Ties matter: Meta-evaluating modern metrics with pairwise accuracy and tie calibration, 2023
Daniel Deutsch, George Foster, and Markus Freitag · 2023
Later among the works it cites.
G-eval: NLG evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu · 2023
Later among the works it cites.
Revisiting the gold standard: Grounding summarization evaluation with robust human evaluation
Yixin Liu, Alex Fabbri, Pengfei Liu, Yilun Zhao, Linyong Nan, Ruilin Han, Simeng Han, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, and Dragomir Radev · 2023
Later among the works it cites.
DKPro agreement: An open-source Java library for measuring inter-rater agreement
Christian M. Meyer, Margot Mieskes, Christian Stab, and Iryna Gurevych · 2023
Later among the works it cites.
Collective Human Opinions in Semantic Textual Similarity
Yuxia Wang, Shimin Tao, Ning Xie, Hao Yang, Timothy Baldwin, and Karin Verspoor · 2023
Later among the works it cites.
AlignScore: Evaluating factual consistency with a unified alignment function
Yuheng Zha, Yichi Yang, Ruichen Li, and Zhiting Hu · 2023
Later among the works it cites.
Judging LLM-as-a-judge with MT-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica · 2023
Later among the works it cites.
Llms instead of human judges? a large scale empirical study across 20 nlp evaluation tasks
Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, Raquel Fernández, Albert Gatt, E. Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, André F. T. Martins, Philipp Mondorf, Vera Neplenbroek, Sandro Pezzelle, Barbara Plank, David Schlangen, Alessandro Suglia, Aditya K Surikuchi, Ece Takmaz, and Alberto Testoni · 2024
Closest in time.
Beiduo Chen, Xinpeng Wang, Siyao Peng, Robert Litschko, Anna Korhonen, and Barbara Plank · 2024
Closest in time.
ConSiDERS-the-human evaluation framework: Rethinking human evaluation for generative large language models
Aparna Elangovan, Ling Liu, Lei Xu, Sravan Babu Bodapati, and Dan Roth · 2024
Closest in time.
Prometheus: Inducing fine-grained evaluation capability in language models
Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, and Minjoon Seo · 2024
Closest in time.