Fetching the paper…
Reading the bibliography…
Given the rapid progress of generative AI, there is a pressing need to systematically compare and choose between the numerous models and configurations available.
A technique for the measurement of attitudes
Rensis Likert. 1932 · 1932
Earlier work this paper cites.
Comparing individual means in the analysis of variance
John W Tukey. 1949 · 1949
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry. 1952 · 1952
Earlier work this paper cites.
An investigation into the validity of some metrics for automatically evaluating natural language generation systems
Ehud Reiter and Anja Belz. 2009 · 2009
Earlier work this paper cites.
Automatically assessing machine summary content without a gold standard
Annie Louis and Ani Nenkova. 2013 · 2013
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
Beta calibration: a well-founded and easily implemented improvement on logistic calibration for binary classifiers
Meelis Kull, Telmo Silva Filho, and Peter Flach. 2017 · 2017
Earlier work this paper cites.
Better than average: Paired evaluation of NLP systems
Maxime Peyrard, Wei Zhao, Steffen Eger, and Robert West. 2021 · 2021
Earlier work this paper cites.
Re-examining system-level correlations of automatic summarization evaluation metrics
Daniel Deutsch, Rotem Dror, and Dan Roth. 2022 · 2022
Earlier work this paper cites.
Verbosity bias in preference labeling by large language models
Keita Saito, Akifumi Wachi, Koki Wataoka, and Youhei Akimoto. 2023 · 2023
Earlier work this paper cites.
Classifier calibration: a survey on how to assess and improve predicted class probabilities
Telmo Silva Filho, Hao Song, Miquel Perello-Nieto, Raul Santos-Rodriguez, Meelis Kull, and Peter Flach. 2023 · 2023
Earlier work this paper cites.
Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback
Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher Manning. 2023 · 2023
Earlier work this paper cites.
Large language models are not fair evaluators
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. 2023 · 2023
Earlier work this paper cites.
Judging LLM-as-a-judge with MT-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E Gonzalez, and Ion Stoica. 2023 · 2023
Earlier work this paper cites.
Label-efficient model selection for text generation
Shir Ashury Tahan, Ariel Gera, Benjamin Sznajder, Leshem Choshen, Liat Ein-Dor, and Eyal Shnarch. 2024 · 2024
Earlier work this paper cites.
LLMs instead of human judges? a large scale empirical study across 20 NLP evaluation tasks
Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, Raquel Fernández, Albert Gatt, Esam Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, et al. 2024 · 2024
Cited alongside, same era.
Zheng Cai, Maosong Cao, Haojiong Chen, Kai Chen, Keyu Chen, Xin Chen, Xun Chen, Zehui Chen, Zhi Chen, Pei Chu, et al. 2024 · 2024
Cited alongside, same era.
ODIN: Disentangled reward mitigates hacking in RLHF
Lichang Chen, Chen Zhu, Jiuhai Chen, Davit Soselia, Tianyi Zhou, Tom Goldstein, Heng Huang, Mohammad Shoeybi, and Bryan Catanzaro. 2024 · 2024
Cited alongside, same era.
Chatbot Arena: An open platform for evaluating LLMs by human preference
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael Jordan, Joseph E Gonzalez, et al. 2024 · 2024
Cited alongside, same era.
Social choice should guide ai alignment in dealing with diverse human feedback
State of what art? a call for multi-prompt LLM evaluation
Moran Mizrahi, Guy Kaplan, Dan Malkin, Rotem Dror, Dafna Shahaf, and Gabriel Stanovsky. 2024 · 2024
Closest in time.
Llm evaluators recognize and favor their own generations
Arjun Panickssery, Samuel R. Bowman, and Shi Feng. 2024 · 2024
Closest in time.
OffsetBias: Leveraging debiased data for tuning evaluators
Junsoo Park, Seungyeon Jwa, Ren Meiying, Daeyoung Kim, and Sanghyuk Choi. 2024 · 2024
Closest in time.
Do these LLM benchmarks agree? Fixing benchmark evaluation with BenchBench
Yotam Perlitz, Ariel Gera, Ofir Arviv, Asaf Yehudai, Elron Bandel, Eyal Shnarch, Michal Shmueli-Scheuer, and Leshem Choshen. 2024 · 2024
Closest in time.
JudgeBench: A benchmark for evaluating LLM-based judges
Sijun Tan, Siyuan Zhuang, Kyle Montgomery, William Y Tang, Alejandro Cuadron, Chenguang Wang, Raluca Ada Popa, and Ion Stoica. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vincent Conitzer, Rachel Freedman, Jobst Heitzig, Wesley H Holliday, Bob M Jacobs, Nathan Lambert, Milan Mossé, Eric Pacuit, Stuart Russell, Hailey Schoelkopf, et al. 2024 · 2024
Cited alongside, same era.
Limits to scalable evaluation at the frontier: LLM as judge won’t beat twice the data
Florian E Dorner, Vivian Y Nastl, and Moritz Hardt. 2024 · 2024
Cited alongside, same era.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024 · 2024
Cited alongside, same era.
Length-controlled AlpacaEval: A simple way to debias automatic evaluators
Yann Dubois, Balázs Galambosi, Percy Liang, and Tatsunori B Hashimoto. 2024 · 2024
Cited alongside, same era.
Style outweighs substance: Failure modes of LLM judges in alignment benchmarking
Benjamin Feuer, Micah Goldblum, Teresa Datta, Sanjana Nambiar, Raz Besaleli, Samuel Dooley, Max Cembalest, and John P Dickerson. 2024 · 2024
Cited alongside, same era.
M-RewardBench: Evaluating reward models in multilingual settings
Srishti Gureja, Lester James Validad Miranda, Shayekh Bin Islam, Rishabh Maheshwary, Drishti Sharma, Gusti Winata, Nathan Lambert, Sebastian Ruder, Sara Hooker, and Marzieh Fadaee. 2024 · 2024
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024 · 2024
Cited alongside, same era.
Hannah Rose Kirk, Alexander Whitefield, Paul Röttger, Andrew Bean, Katerina Margatina, Juan Ciro, Rafael Mosquera, Max Bartolo, Adina Williams, He He, Bertie Vidgen, and Scott A Hale. 2024 · 2024
Cited alongside, same era.
Closest in time.
Judging the judges: Evaluating alignment and vulnerabilities in LLMs-as-judges
Aman Singh Thakur, Kartik Choudhary, Venkat Srinik Ramayapally, Sankaran Vaidyanathan, and Dieuwke Hupkes. 2024 · 2024
Closest in time.
A measure of the system dependence of automated metrics
Pius von Däniken, Jan Deriu, and Mark Cieliebak. 2024 · 2024
Closest in time.
Favi-Score: A measure for favoritism in automated preference ratings for generative AI evaluation
Pius Von Däniken, Jan Deriu, Don Tuggener, and Mark Cieliebak. 2024 · 2024
Closest in time.
Interpretable preferences via multi-objective reward modeling and mixture-of-experts
Haoxiang Wang, Wei Xiong, Tengyang Xie, Han Zhao, and Tong Zhang. 2024 · 2024
Closest in time.
Hui Wei, Shenghua He, Tian Xia, Andy Wong, Jingyang Lin, and Mei Han. 2024 · 2024
Closest in time.
Pride and prejudice: LLM amplifies self-bias in self-refinement
Wenda Xu, Guanglei Zhu, Xuandong Zhao, Liangming Pan, Lei Li, and William Wang. 2024 · 2024
Closest in time.
Regularizing hidden states enables learning generalizable reward model for LLMs
Rui Yang, Ruomeng Ding, Yong Lin, Huan Zhang, and Tong Zhang. 2024 · 2024
Closest in time.
Justice or prejudice? quantifying biases in LLM-as-a-judge
Jiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen, Qihui Zhang, Nuno Moniz, Tian Gao, Werner Geyer, Chao Huang, Pin-Yu Chen, et al. 2024 · 2024
Closest in time.
Achieving human parity in content-grounded datasets generation
Asaf Yehudai, Boaz Carmeli, Yosi Mass, Ofir Arviv, Nathaniel Mills, Eyal Shnarch, and Leshem Choshen. 2024 · 2024
Closest in time.
Advancing LLM reasoning generalists with preference trees
Lifan Yuan, Ganqu Cui, Hanbin Wang, Ning Ding, Xingyao Wang, Jia Deng, Boji Shan, Huimin Chen, Ruobing Xie, Yankai Lin, Zhenghao Liu, Bowen Zhou, Hao Peng, Zhiyuan Liu, and Maosong Sun. 2024 · 2024
Closest in time.