Fetching the paper…
Reading the bibliography…
Automatic evaluation methods based on large language models (LLMs) are emerging as the standard tool for assessing the instruction-following abilities of LLM-based agents.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
The USCF Rating System: Its Development, Theory, and Applications
Elo, A · 1966
Earlier work this paper cites.
A generalized model for multidimensional intransitivity
Duan, J., Li, J., Baba, Y., and Kashima, H · 2017
Earlier work this paper cites.
Open-ended learning in symmetric zero-sum games
Balduzzi, D., Garnelo, M., Bachrach, Y., Czarnecki, W., Perolat, J., Jaderberg, M., and Graepel, T · 2019
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning, 2019
OpenAI, :, Berner, C., Brockman, G., Chan, B., Cheung, V., Dębiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., Józefowicz, R., Gray, S., Olsson, C., Pachocki, J., Petrov, M., d. O. Pinto, H. P., Raiman, J., Salimans, T., Schlatter, J., Schneider, J., Sidor, S., Sutskever, I., Tang, J., Wolski, F., and Zhang, S · 2019
Earlier work this paper cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W., Mathieu, M., Dudzik, A., Chung, J., Choi, D., Powell, R., Ewalds, T., Georgiev, P., Oh, J., Horgan, D., Kroiss, M., Danihelka, I., Huang, A., Sifre, L., Cai, T., Agapiou, J., Jaderberg, M., and Silver, D · 2019
Earlier work this paper cites.
Real world games look like spinning tops
Czarnecki, W. M., Gidel, G., Tracey, B., Tuyls, K., Omidshafiei, S., Balduzzi, D., and Jaderberg, M · 2020
Earlier work this paper cites.
The perils of using Mechanical Turk to evaluate open-ended text generation
Karpinska, M., Akoury, N., and Iyyer, M · 2021
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., and Lowe, R · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., ichter, b., Xia, F., Chi, E., Le, Q. V., and Zhou, D · 2022
Earlier work this paper cites.
Model card and evaluations for claude models, 2023
Anthropic · 2023
Earlier work this paper cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Earlier work this paper cites.
Gemini: A family of highly capable multimodal models
Gemini Team Google · 2023
Earlier work this paper cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Earlier work this paper cites.
Alpacaeval: An automatic evaluator of instruction-following models, 5 2023
Li, X., Zhang, T., Dubois, Y., Taori, R., Gulrajani, I., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Cited alongside, same era.
Gpt-4 technical report, 2023
OpenAI, Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., and othres · 2023
Cited alongside, same era.
Verbosity bias in preference labeling by large language models, 2023
Saito, K., Wachi, A., Wataoka, K., and Akimoto, Y · 2023
Cited alongside, same era.
Judging LLM-as-a-judge with MT-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., Zhang, H., Gonzalez, J. E., and Stoica, I · 2023
Cited alongside, same era.
Yi-large llm launch, 2024
01.AI · 2024
Cited alongside, same era.
Yi: Open foundation models by 01.ai, 2024
01.AI, Young, A., Chen, B., Li, C., Huang, C., Zhang, G., Zhang, G., Li, H., Zhu, J., Chen, J., Chang, J., Yu, K., Liu, P., Liu, Q., Yue, S., Yang, S., Yang, S., Yu, T., Xie, W., Huang, W., Hu, X., Ren, X., Niu, X., Nie, P., Xu, Y., Liu, Y., Wang, Y., Cai, Y., Gu, Z., Liu, Z., and Dai, Z · 2024
From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline, 2024
Li, T., Chiang, W.-L., Frick, E., Dunlap, L., Wu, T., Zhu, B., Gonzalez, J. E., and Stoica, I · 2024
Later among the works it cites.
Aligning with human judgement: The role of pairwise preference in large language model evaluators
Liu, Y., Zhou, H., Guo, Z., Shareghi, E., Vulić, I., Korhonen, A., and Collier, N · 2024
Later among the works it cites.
LLM comparative assessment: Zero-shot NLG evaluation through pairwise comparisons using large language models
Liusie, A., Manakul, P., and Gales, M. J. F · 2024
Later among the works it cites.
Large language models sensitivity to the order of options in multiple-choice questions
Pezeshkpour, P. and Hruschka, E · 2024
Later among the works it cites.
Introducing qwen1.5, 2024
Qwen Team · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
When benchmarks are targets: Revealing the sensitivity of large language model leaderboards
Alzahrani, N., Alyahya, H., Alnumay, Y., AlRashed, S., Alsubaie, S., Almushayqih, Y., Mirza, F., Alotaibi, N., Al-Twairesh, N., Alowisheq, A., Bari, M. S., and Khan, H · 2024
Cited alongside, same era.
The claude 3 model family: Opus, sonnet, haiku, 2024
Anthropic, A · 2024
Cited alongside, same era.
Humans or LLMs as the judge? a study on judgement bias
Chen, G. H., Chen, S., Liu, Z., Jiang, F., and Wang, B · 2024
Cited alongside, same era.
Chatbot arena: An open platform for evaluating LLMs by human preference
Chiang, W.-L., Zheng, L., Sheng, Y., Angelopoulos, A. N., Li, T., Li, D., Zhu, B., Zhang, H., Jordan, M., Gonzalez, J. E., and Stoica, I · 2024
Cited alongside, same era.
Ticking all the boxes: Generated checklists improve llm evaluation and generation, 2024
Cook, J., Rocktäschel, T., Foerster, J., Aumiller, D., and Wang, A · 2024
Cited alongside, same era.
Length-controlled alpacaeval: A simple debiasing of automatic evaluators
Dubois, Y., Liang, P., and Hashimoto, T · 2024
Cited alongside, same era.
Raina, V., Liusie, A., and Gales, M · 2024
Later among the works it cites.
Rainbow teaming: Open-ended generation of diverse adversarial prompts
Samvelyan, M., Raparthy, S. C., Lupu, A., Hambro, E., Markosyan, A. H., Bhatt, M., Mao, Y., Jiang, M., Parker-Holder, J., Foerster, J. N., Rocktäschel, T., and Raileanu, R · 2024
Later among the works it cites.
Large language models are not fair evaluators
Wang, P., Li, L., Chen, L., Cai, Z., Zhu, D., Lin, B., Cao, Y., Kong, L., Liu, Q., Liu, T., and Sui, Z · 2024
Later among the works it cites.
WizardLM: Empowering large pre-trained language models to follow complex instructions
Xu, C., Sun, Q., Zheng, K., Geng, X., Zhao, P., Feng, J., Tao, C., Lin, Q., and Jiang, D · 2024
Later among the works it cites.
Fairer preferences elicit improved human-aligned large language model judgments
Zhou, H., Wan, X., Liu, Y., Collier, N., Vulić, I., and Korhonen, A · 2024
Later among the works it cites.
Starling-7b: Improving helpfulness and harmlessness with RLAIF
Zhu, B., Frick, E., Wu, T., Zhu, H., Ganesan, K., Chiang, W.-L., Zhang, J., and Jiao, J · 2024
Later among the works it cites.
Wildbench: Benchmarking LLMs with challenging tasks from real users in the wild
Lin, B. Y., Deng, Y., Chandu, K., Ravichander, A., Pyatkin, V., Dziri, N., Bras, R. L., and Choi, Y · 2025
Closest in time.
Livebench: A challenging, contamination-limited LLM benchmark
White, C., Dooley, S., Roberts, M., Pal, A., Feuer, B., Jain, S., Shwartz-Ziv, R., Jain, N., Saifullah, K., Dey, S., Shubh-Agrawal, Sandha, S. S., Naidu, S. V., Hegde, C., LeCun, Y., Goldstein, T., Neiswanger, W., and Goldblum, M · 2025
Closest in time.