Fetching the paper…
Reading the bibliography…
LLM-based auto-annotators have become a key component of the LLM development process due to their cost-effectiveness and scalability compared to human-based evaluation.
The rating of chessplayers: Past and present
Arpad E Elo · 1978
Earlier work this paper cites.
Automatically assessing machine summary content without a gold standard
Annie Louis and Ani Nenkova · 2002
Earlier work this paper cites.
Causality: Models, Reasoning and Inference
Judea Pearl · 2009
Earlier work this paper cites.
Causal inference, 2010
Miguel A Hernán and James M Robins · 2010
Earlier work this paper cites.
Controlled direct and mediated effects: Definition, identification and bounds
Tyler J. VanderWeele · 2010
Earlier work this paper cites.
Explanation in causal inference: methods for mediation and interaction
Tyler VanderWeele · 2015
Earlier work this paper cites.
Why we need new evaluation metrics for NLG
J. Novikova, O. Dušek, A. C. Curry, and V. Rieser · 2017
Earlier work this paper cites.
Evaluating factuality in generation with dependency-level entailment
Tanya Goyal and Greg Durrett · 2020
Earlier work this paper cites.
Evaluating the factual consistency of abstractive text summarization
Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher · 2020
Earlier work this paper cites.
Learning an unreferenced metric for online dialogue evaluation
Koustuv Sinha, Prasanna Parthasarathi, Jasmine Wang, Ryan Lowe, William L. Hamilton, and Joelle Pineau · 2020
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi · 2020
Cited alongside, same era.
A comprehensive assessment of dialog evaluation metrics
Y. Yeh, M. Eskenazi, and S. Mehri · 2021
Cited alongside, same era.
Spurious correlations in reference-free evaluation of text generation
Esin Durmus, Faisal Ladhak, and Tatsunori Hashimoto · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Alpacafarm: A simulation framework for methods that learn from human feedback, 2023
Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Cited alongside, same era.
Large language models are not fair evaluators
Peiyi Wang, Lei Li, Liang Chen, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui · 2023
Later among the works it cites.
Style over substance: Evaluation biases for large language models
Minghao Wu and Alham Fikri Aji · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Haotong Zhang, Joseph Gonzalez, and Ion Stoica · 2023
Later among the works it cites.
Odin: Disentangled reward mitigates hacking in rlhf
Lichang Chen, Chen Zhu, Davit Soselia, Jiuhai Chen, Tianyi Zhou, Tom Goldstein, Heng Huang, Mohammad Shoeybi, and Bryan Catanzaro · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang · 2023
Cited alongside, same era.
Alpacaeval: An automatic evaluator of instruction-following models
Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Cited alongside, same era.
Loose lips sink ships: Mitigating length bias in reinforcement learning from human feedback
Wei Shen, Rui Zheng, Wenyu Zhan, Jun Zhao, Shihan Dou, Tao Gui, Qi Zhang, and Xuanjing Huang · 2023
Cited alongside, same era.
A long way to go: Investigating length correlations in rlhf, 2023
Prasann Singhal, Tanya Goyal, Jiacheng Xu, and Greg Durrett · 2023
Cited alongside, same era.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Cited alongside, same era.
Viet Hoang Tran Duong · 2024
Closest in time.
Advanced length-normalized alpacaeval 2.0, 2024
Balazs Galambosi · 2024
Closest in time.
Wildbench: Benchmarking llms with challenging tasks from real users in the wild, 2024
Bill Yuchen Lin, Khyathi Chandu, Faeze Brahman, Yuntian Deng, Abhilasha Ravichander, Valentina Pyatkin, Ronan Le Bras, and Yejin Choi · 2024
Closest in time.
Disentangling length from quality in direct preference optimization
Ryan Park, Rafael Rafailov, Stefano Ermon, and Chelsea Finn · 2024
Closest in time.
Length-normalized alpacaeval 2.0, 2024
Teortaxes · 2024
Closest in time.