Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are increasingly deployed as automatic judges to evaluate system outputs in tasks such as summarization, dialogue, and creative writing.
ELI5: Long form question answering
A. Fan, Y. Jernite, E. Perez, D. Grangier, J. Weston, and M. Auli · 2019
Earlier work this paper cites.
Summeval: Re-evaluating summarization evaluation
A. R. Fabbri, W. Krysciński, B. McCann, C. Xiong, R. Socher, and D. Radev · 2021
Earlier work this paper cites.
Human-bot comparison as an evaluation framework for dialogue
S. Mehri and M. Eskenazi · 2022
Earlier work this paper cites.
G-eval: General evaluation of language models
X. Liu et al · 2023
Earlier work this paper cites.
Verbosity bias in preference labeling by large language models
K. Saito, A. Wachi, K. Wataoka, and Y. Akimoto · 2023
Earlier work this paper cites.
M. Turpin, J. Michael, E. Perez, and S. R. Bowman · 2023
Cited alongside, same era.
Chatbot arena: A human-ai comparative judgment platform
L. Zheng et al · 2023
Cited alongside, same era.
Length-controlled alpacaeval: A simple way to debias automatic evaluators
Y. Dubois, B. Galambosi, P. Liang, and T. B. Hashimoto · 2024
Cited alongside, same era.
Llm evaluators recognize and favor their own generations
A. Panickssery, S. R. Bowman, and S. Feng · 2024
Cited alongside, same era.
Judging the judges: A systematic study of position bias in llm-as-a-judge
L. Shi, C. Ma, W. Liang, W. Ma, and S. Vosoughi · 2024
Cited alongside, same era.
Judging LLM/as-a-judge with MT-bench and chatbot arena
L. Zheng, W. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Li, D. Li, E. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica
Cited in the paper.
Self-preference bias in llm-as-a-judge
K. Wataoka, T. Takahashi, and R. Ri · 2024
Later among the works it cites.
Chain-of-thought reasoning in the wild is not always faithful
I. Arcuschin, J. Janiak, R. Krzyzanowski, S. Rajamanoharan, N. Nanda, and A. Conmy · 2025
Closest in time.
Reasoning models don’t always say what they think
Y. Chen, J. Benton, A. Radhakrishnan, J. Uesato, C. Denison, J. Schulman, A. Somani, P. Hase, M. Wagner, F. Roger, et al · 2025
Closest in time.
LitBench: A benchmark and dataset for reliable evaluation of creative writing
D. Fein, S. Russo, V. Xiang, K. Jolly, R. Rafailov, and N. Haber · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…