Fetching the paper…
Reading the bibliography…
Open community-driven platforms like Chatbot Arena that collect user preference data from site visitors have gained a reputation as one of the most trustworthy publicly available benchmarks for LLM performance.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry. 1952 · 1952
Earlier work this paper cites.
Captcha: Using hard ai problems for security
Luis Von Ahn, Manuel Blum, Nicholas J Hopper, and John Langford. 2003 · 2003
Earlier work this paper cites.
Evaluation of text generation: A survey
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020 · 2006
Earlier work this paper cites.
A content-driven reputation system for the wikipedia
B. Thomas Adler and Luca de Alfaro. 2007 · 2007
Earlier work this paper cites.
recaptcha: Human-based character recognition via web security measures
Luis Von Ahn, Benjamin Maurer, Colin McMillen, David Abraham, and Manuel Blum. 2008 · 2008
Earlier work this paper cites.
Accurately detecting trolls in slashdot zoo via decluttering
Srijan Kumar, Francesca Spezzano, and VS Subrahmanian. 2014 · 2014
Earlier work this paper cites.
Why is that relevant? collecting annotator rationales for relevance judgments
Tyler McDonnell, Matthew Lease, Mucahid Kutlu, and Tamer Elsayed. 2016 · 2016
Earlier work this paper cites.
The troll-trust model for ranking in signed networks
Zhaoming Wu, Charu C Aggarwal, and Jimeng Sun. 2016 · 2016
Earlier work this paper cites.
Your behavior signals your reliability: Modeling crowd behavioral traces to ensure quality relevance annotations
Tanya Goyal, Tyler McDonnell, Mucahid Kutlu, Tamer Elsayed, and Matthew Lease. 2018 · 2018
Earlier work this paper cites.
All that’s ‘human’is not gold: Evaluating human evaluation of generated text
Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A Smith. 2021 · 2021
Cited alongside, same era.
The perils of using mechanical turk to evaluate open-ended text generation
Marzena Karpinska, Nader Akoury, and Mohit Iyyer. 2021 · 2021
Cited alongside, same era.
Evaluation examples are not equally informative: How should that change nlp leaderboards?
Pedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P Lalor, Robin Jia, and Jordan Boyd-Graber. 2021 · 2021
Cited alongside, same era.
Snac: Coherence error detection for narrative summarization
Tanya Goyal, Junyi Jessy Li, and Greg Durrett. 2022b · 2022
Cited alongside, same era.
Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text
Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam. 2023 · 2023
Cited alongside, same era.
Human feedback is not gold standard
Tom Hosking, Phil Blunsom, and Max Bartolo. 2024 · 2024
Closest in time.
Rewardbench: Evaluating reward models for language modeling
Nathan Lambert, Valentina Pyatkin, Jacob Morrison, LJ Miranda, Bill Yuchen Lin, Khyathi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, et al. 2024 · 2024
Closest in time.
From live data to high-quality benchmarks: The arena-hard pipeline
Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Banghua Zhu, Joseph E. Gonzalez, and Ion Stoica. 2024 · 2024
Closest in time.
Wildbench: Benchmarking llms with challenging tasks from real users in the wild
Bill Yuchen Lin, Yuntian Deng, Khyathi Chandu, Faeze Brahman, Abhilasha Ravichander, Valentina Pyatkin, Nouha Dziri, Ronan Le Bras, and Yejin Choi. 2024 · 2024
Closest in time.
Wildvision: Evaluating vision-language models in the wild with human preferences
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A watermark for large language models
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein. 2023 · 2023
Cited alongside, same era.
Longeval: Guidelines for human evaluation of faithfulness in long-form summarization
Kalpesh Krishna, Erin Bransom, Bailey Kuehl, Mohit Iyyer, Pradeep Dasigi, Arman Cohan, and Kyle Lo. 2023 · 2023
Cited alongside, same era.
Alpacaeval: An automatic evaluator of instruction-following models
Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023 · 2023
Cited alongside, same era.
Chatbot arena: An open platform for evaluating llms by human preference
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica. 2024a
Cited in the paper.
Chatbot arena: An open platform for evaluating llms by human preference
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael Jordan, Joseph E Gonzalez, et al. 2024b
Cited in the paper.
A case for better evaluation standards in nlg
Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam
Cited in the paper.
Yujie Lu, Dongfu Jiang, Wenhu Chen, William Yang Wang, Yejin Choi, and Bill Yuchen Lin. 2024 · 2024
Closest in time.
Contextualized evaluations: Taking the guesswork out of language model evaluations
Chaitanya Malaviya, Joseph Chee Chang, Dan Roth, Mohit Iyyer, Mark Yatskar, and Kyle Lo. 2024 · 2024
Closest in time.
Mixeval: Deriving wisdom of the crowd from llm benchmark mixtures
Jinjie Ni, Fuzhao Xue, Xiang Yue, Yuntian Deng, Mahir Shah, Kabir Jain, Graham Neubig, and Yang You. 2024 · 2024
Closest in time.
Researchy questions: A dataset of multi-perspective, decompositional questions for llm web agents
Corby Rosset, Ho-Lam Chung, Guanghui Qin, Ethan C Chau, Zhuo Feng, Ahmed Awadallah, Jennifer Neville, and Nikhil Rao. 2024 · 2024
Closest in time.