Fetching the paper…
Reading the bibliography…
Arena-based evaluation is a fundamental yet significant evaluation paradigm for modern AI models, especially large language models (LLMs).
Die berechnung der turnier-ergebnisse als ein maximumproblem der wahrscheinlichkeitsrechnung
Zermelo, E · 1929
Earlier work this paper cites.
The proposed uscf rating system, its development, theory, and applications
Elo, A. E · 1967
Earlier work this paper cites.
Ties in paired-comparison experiments: A generalization of the bradley-terry model
Rao, P. and Kupper, L. L · 1967
Earlier work this paper cites.
Positive definite matrices
Johnson, C. R · 1970
Earlier work this paper cites.
The role of the hessian matrix in fitting models to measurements
Thacker, W. C · 1989
Earlier work this paper cites.
Sequential design of experiments
Chernoff, H · 1992
Earlier work this paper cites.
Classification by pairwise coupling
Hastie, T. and Tibshirani, R · 1997
Earlier work this paper cites.
The theory of newton’s method
Galántai, A · 2000
Earlier work this paper cites.
Solving nonlinear equations with Newton’s method
Kelley, C. T · 2003
Earlier work this paper cites.
Convex optimization
Boyd, S. and Vandenberghe, L · 2004
Earlier work this paper cites.
Mm algorithms for generalized bradley-terry models
Hunter, D. R · 2004
Earlier work this paper cites.
Combining svms with various feature selection strategies
Chen, Y.-W. and Lin, C.-J · 2006
Earlier work this paper cites.
Computing “elo ratings” of move patterns in the game of go
Coulom, R · 2007
Earlier work this paper cites.
Toward modern psychometrics
Morizot, J., Ainsworth, A. T., and Reise, S. P · 2009
Earlier work this paper cites.
How reliable are annotations via crowdsourcing: a study about inter-annotator agreement for multi-label image annotation
Nowak, S. and Rüger, S · 2010
Earlier work this paper cites.
How i won the” chess ratings-elo vs the rest of the world” competition
Sismanis, Y · 2010
Earlier work this paper cites.
Online crowdsourcing: rating annotators and obtaining cost-effective labels
Welinder, P. and Perona, P · 2010
Earlier work this paper cites.
Ranking annotators for crowdsourced labeling tasks
Raykar, V. C. and Yu, S · 2011
Earlier work this paper cites.
Item response theory
Embretson, S. E. and Reise, S. P · 2013
Earlier work this paper cites.
Introduction to quadratic forms , volume 117
O’Meara, O. T · 2013
Cited alongside, same era.
Preference-based rank elicitation using statistical models: The case of mallows
Busa-Fekete, R., Hüllermeier, E., and Szörényi, B · 2014
Cited alongside, same era.
Application of lagrange mean value theorem
Shi-gu, J · 2014
Cited alongside, same era.
Online rank elicitation for plackett-luce: A dueling bandits approach
Szörényi, B., Busa-Fekete, R., Paul, A., and Hüllermeier, E · 2015
Cited alongside, same era.
Applications of the elo rating system in adaptive educational systems
Pelánek, R · 2016
Cited alongside, same era.
An overview of gradient descent optimization algorithms
Ruder, S · 2016
Fully adaptive framework: Neural computerized adaptive testing for online education
Zhuang, Y., Liu, Q., Huang, Z., Li, Z., Shen, S., and Ma, H · 2022
Later among the works it cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Later among the works it cites.
Elo uncovered: Robustness and best practices in language model evaluation
Boubdir, M., Kim, E., Ermis, B., Hooker, S., and Fadaee, M · 2023
Later among the works it cites.
Generative judge for evaluating alignment
Li, J., Sun, S., Yuan, W., Fan, R.-Z., Zhao, H., and Liu, P · 2023
Later among the works it cites.
Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Elo ratings and the sports model: A neglected topic in applied probability?
Aldous, D · 2017
Cited alongside, same era.
Cognitive biases in crowdsourcing
Eickhoff, C · 2018
Cited alongside, same era.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Cited alongside, same era.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Cited alongside, same era.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Cited alongside, same era.
An elo-like system for massive multiplayer competitions
Ebtekar, A. and Liu, P · 2021
Cited alongside, same era.
Wang, Y., Yu, Z., Zeng, Z., Yang, L., Wang, C., Chen, H., Jiang, C., Xie, R., Wang, J., Xie, X., et al · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I · 2023
Later among the works it cites.
Judgelm: Fine-tuned large language models are scalable judges
Zhu, L., Wang, X., and Wang, X · 2023
Later among the works it cites.
Llmchain: Blockchain-based reputation system for sharing and evaluating large language models
Bouchiha, M. A., Telnoff, Q., Bakkali, S., Champagnat, R., Rabah, M., Coustaty, M., and Ghamri-Doudane, Y · 2024
Later among the works it cites.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chan, C., Chen, W., Su, Y., Yu, J., Xue, W., Zhang, S., Fu, J., and Liu, Z · 2024
Later among the works it cites.
Towards personalized evaluation of large language models with an anonymous crowd-sourcing platform
Cheng, M., Zhang, H., Yang, J., Liu, Q., Li, L., Huang, X., Song, L., Li, Z., Huang, Z., and Chen, E · 2024
Later among the works it cites.
Chatbot arena: An open platform for evaluating llms by human preference
Chiang, W., Zheng, L., Sheng, Y., Angelopoulos, A. N., Li, T., Li, D., Zhu, B., Zhang, H., Jordan, M. I., Gonzalez, J. E., and Stoica, I · 2024
Later among the works it cites.
A merge sort based ranking system for the evaluation of large language models
Li, C., Shi, L., Zhou, C., Huan, Z., Tang, C., Zhang, X., Wang, X., Zhou, J., and Liu, S · 2024
Later among the works it cites.
Computerized adaptive testing via collaborative ranking
Liu, Z., Yan, Z., Liu, Q., Li, J., Zhang, Y., Huang, Z., Wu, J., and Wang, S · 2024
Later among the works it cites.
tinybenchmarks: evaluating llms with fewer examples
Polo, F. M., Weber, L., Choshen, L., Sun, Y., Xu, G., and Yurochkin, M · 2024
Later among the works it cites.
Evaluatology: The science and engineering of evaluation
Zhan, J., Wang, L., Gao, W., Li, H., Wang, C., Huang, Y., Li, Y., Yang, Z., Kang, G., Luo, C., et al · 2024
Later among the works it cites.
Competeai: Understanding the competition behaviors in large language model-based agents
Zhao, Q., Wang, J., Zhang, Y., Jin, Y., Zhu, K., Chen, H., and Xie, X · 2024
Later among the works it cites.
A survey on knowledge-oriented retrieval-augmented generation
Cheng, M., Luo, Y., Ouyang, J., Liu, Q., Liu, H., Li, L., Yu, S., Zhang, B., Cao, J., Ma, J., et al · 2025
Closest in time.
Ouyang, J., Pan, T., Cheng, M., Yan, R., Luo, Y., Lin, J., and Liu, Q · 2025
Closest in time.