Fetching the paper…
Reading the bibliography…
Chatbot Arena is a popular platform for evaluating LLMs by pairwise battles, where users vote for their preferred response from two randomly sampled anonymous models.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
The proposed uscf rating system, its development, theory, and applications
Elo, A. E · 1967
Earlier work this paper cites.
Manipulation of voting schemes: a general result
Gibbard, A · 1973
Earlier work this paper cites.
The vulnerability of point-voting schemes to preference variation and strategic manipulation
Nitzan, S · 1985
Earlier work this paper cites.
The manipulation of voting systems
Hartvigsen, D · 2008
Earlier work this paper cites.
Voting systems with trust mechanisms in cyberspace: Vulnerabilities and defenses
Feng, Q., Sun, Y. L., Liu, L., Yang, Y., and Dai, Y · 2010
Earlier work this paper cites.
Methods of voting system and manipulation of voting
Islam, J., Mohajan, H., and Moolio, P · 2010
Earlier work this paper cites.
Voting systems and strategic manipulation: An experimental study
Bassi, A · 2015
Earlier work this paper cites.
Using information theory to improve the robustness of trust systems
Wang, D., Muller, T., Irissappane, A. A., Zhang, J., and Liu, Y · 2015
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2018
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. D. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
Is gpt-3 text indistinguishable from human text? scarecrow: A framework for scrutinizing machine text
Dou, Y., Forbes, M., Koncel-Kedziorski, R., Smith, N. A., and Choi, Y · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Guiding neural story generation with reader models
Peng, X., Xie, K., Alabdulkarim, A., Kayam, H., Dani, S., and Riedl, M · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., Huang, F., et al · 2023
Earlier work this paper cites.
Chatbot arena: New models & elo system update, 2023a
Chiang, W.-L., Li, T. L., Gonzalez, J. E., and Stoica, I · 2023
Earlier work this paper cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., et al · 2023
Earlier work this paper cites.
Three bricks to consolidate watermarks for large language models
Fernandez, P., Chaffin, A., Tit, K., Chappelier, V., and Furon, T · 2023
Earlier work this paper cites.
A survey on the possibilities & impossibilities of ai-generated text detection
Ghosal, S. S., Chakraborty, S., Geiping, J., Huang, F., Manocha, D., and Bedi, A · 2023
Earlier work this paper cites.
How close is chatgpt to human experts? comparison corpus, evaluation, and detection
Guo, B., Zhang, X., Wang, Z., Jiang, M., Nie, J., Ding, Y., Yue, J., and Wu, Y · 2023
Cited alongside, same era.
dolphin-2.2.1-mistral-7b, 2023
Hartford, E · 2023
Cited alongside, same era.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Cited alongside, same era.
Solar 10.7 b: Scaling large language models with simple yet effective depth up-scaling
Kim, D., Park, C., Kim, S., Lee, W., Song, W., Kim, Y., Kim, H., Kim, Y., Lee, H., Kim, J., et al · 2023
Cited alongside, same era.
A watermark for large language models
Kirchenbauer, J., Geiping, J., Wen, Y., Katz, J., Miers, I., and Goldstein, T · 2023
Cited alongside, same era.
Chatbot arena: An open platform for evaluating llms by human preference
Chiang, W.-L., Zheng, L., Sheng, Y., Angelopoulos, A. N., Li, T., Li, D., Zhu, B., Zhang, H., Jordan, M., Gonzalez, J. E., et al · 2024
Later among the works it cites.
Undetectable watermarks for language models
Christ, M., Gunn, S., and Zamir, O · 2024
Later among the works it cites.
The command r model, 2024
Cohere · 2024
Later among the works it cites.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Later among the works it cites.
Length-controlled alpacaeval: A simple way to debias automatic evaluators
Dubois, Y., Galambosi, B., Liang, P., and Hashimoto, T. B · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
AlpacaEval: An Automatic Evaluator of Instruction-following Models, May 2023
Li, X., Zhang, T., Dubois, Y., Taori, R., Gulrajani, I., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Cited alongside, same era.
Open-ended long text generation via masked language modeling
Liang, X., Tang, Z., Li, J., and Zhang, M · 2023
Cited alongside, same era.
Introducing mpt-7b: A new standard for open-source, commercially usable llms, 2023
Mosaic · 2023
Cited alongside, same era.
On the safety of group manipulation
Peters, H. and Veselova, Y · 2023
Cited alongside, same era.
Universal jailbreak backdoors from poisoned human feedback
Rando, J. and Tramèr, F · 2023
Cited alongside, same era.
Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark
Sainz, O., Campos, J. A., García-Ferrero, I., Etxaniz, J., de Lacalle, O. L., and Agirre, E · 2023
Cited alongside, same era.
On the exploitability of instruction tuning
Shu, M., Wang, J., Zhu, C., Geiping, J., Xiao, C., and Goldstein, T · 2023
Cited alongside, same era.
GLM, T., Zeng, A., Xu, B., Wang, B., Zhang, C., Yin, D., Zhang, D., Rojas, D., Feng, G., Zhao, H., et al · 2024
Later among the works it cites.
Sleeper agents: Training deceptive llms that persist through safety training
Hubinger, E., Denison, C., Mu, J., Lambert, M., Tong, M., MacDiarmid, M., Lanham, T., Ziegler, D. M., Maxwell, T., Cheng, N., et al · 2024
Later among the works it cites.
On the reliability of watermarks for large language models
Kirchenbauer, J., Geiping, J., Wen, Y., Shu, M., Saifullah, K., Kong, K., Fernando, K., Saha, A., Goldblum, M., and Goldstein, T · 2024
Later among the works it cites.
Does style matter? disentangling style and substance in chatbot arena, 2024a
Li, T., Angelopoulos, A., and Chiang, W.-L · 2024
Later among the works it cites.
Introducing hard prompts category in chatbot arena, 2024b
Li, T., Chiang, W.-L., and Dunlap, L · 2024
Later among the works it cites.
Chatbot arena categories definitions, methods, and insights, 2024c
Li, T., Chiang, W.-L., Song, Y., Jain, N., Dunlap, L., Li, D., Frick, E., and Angelopoulos, A. N · 2024
Later among the works it cites.
Rm-bench: Benchmarking reward models of language models with subtlety and style
Liu, Y., Yao, Z., Min, R., Cao, Y., Hou, L., and Li, J · 2024
Later among the works it cites.
Introducing openai o1 preview, 2024
OpenAI · 2024
Later among the works it cites.
Is llm-as-a-judge robust? investigating universal adversarial attacks on zero-shot llm assessment
Raina, V., Liusie, A., and Gales, M · 2024
Later among the works it cites.
Optimization-based prompt injection attack to llm-as-a-judge
Shi, J., Yuan, Z., Liu, Y., Huang, Y., Zhou, P., Sun, L., and Gong, N. Z · 2024
Later among the works it cites.
Ghostbuster: Detecting text ghostwritten by large language models
Verma, V., Fleisig, E., Tomlin, N., and Klein, D · 2024
Later among the works it cites.
Instructions as backdoors: Backdoor vulnerabilities of instruction tuning for large language models
Xu, J., Ma, M., Wang, F., Xiao, C., and Chen, M · 2024
Later among the works it cites.
Backdooring instruction-tuned large language models with virtual prompt injection
Yan, J., Yadav, V., Li, S., Chen, L., Tang, Z., Wang, H., Srinivasan, V., Ren, X., and Jin, H · 2024
Later among the works it cites.
Yi: Open foundation models by 01. ai
Young, A., Chen, B., Li, C., Huang, C., Zhang, G., Zhang, G., Li, H., Zhu, J., Chen, J., Chang, J., et al · 2024
Later among the works it cites.
Challenges in trustworthy human evaluation of chatbots
Zhao, W., Rush, A. M., and Goyal, T · 2024
Later among the works it cites.
Cheating automatic llm benchmarks: Null models achieve high win rates
Zheng, X., Pang, T., Du, C., Liu, Q., Jiang, J., and Lin, M · 2024
Later among the works it cites.
Starling-7b: Improving helpfulness and harmlessness with rlaif
Zhu, B., Frick, E., Wu, T., Zhu, H., Ganesan, K., Chiang, W.-L., Zhang, J., and Jiao, J · 2024
Later among the works it cites.
Exploring and mitigating adversarial manipulation of voting-based leaderboards
Huang, Y., Nasr, M., Angelopoulos, A., Carlini, N., Chiang, W.-L., Choquette-Choo, C. A., Ippolito, D., Jagielski, M., Lee, K., Liu, K. Z., et al · 2025
Closest in time.