Fetching the paper…
Reading the bibliography…
Evaluation of large language model (LLM) outputs requires users to make critical judgments about the best outputs across various configurations.
ACUTE-EVAL: Improved Dialogue Evaluation with Optimized Questions and Multi-turn Comparisons
Margaret Li, Jason Weston, and Stephen Roller. 2019 · 1909
Earlier work this paper cites.
NASA-task load index (NASA-TLX); 20 years later
Sandra G Hart. 2006 · 2006
Earlier work this paper cites.
A general inductive approach for analyzing qualitative evaluation data
David R Thomas. 2006 · 2006
Earlier work this paper cites.
Accounting for diversity in subjective judgments
Evangelos Karapanos, Jean-Bernard Martens, and Marc Hassenzahl. 2009 · 2009
Earlier work this paper cites.
Gestalt: integrated support for implementation and analysis in machine learning
Kayur Patel, Naomi Bancroft, Steven M Drucker, James Fogarty, Amy J Ko, and James Landay. 2010 · 2010
Earlier work this paper cites.
Bigbench: Towards an industry standard benchmark for big data analytics
Ahmad Ghazal, Tilmann Rabl, Minqing Hu, Francois Raab, Meikel Poess, Alain Crolotte, and Hans-Arno Jacobsen. 2013 · 2013
Earlier work this paper cites.
The future of crowd work
Aniket Kittur, Jeffrey V Nickerson, Michael Bernstein, Elizabeth Gerber, Aaron Shaw, John Zimmerman, Matt Lease, and John Horton. 2013 · 2013
Earlier work this paper cites.
Ice: enabling non-experts to build models interactively for large-scale lopsided problems
Patrice Simard, David Chickering, Aparna Lakshmiratan, Denis Charles, Léon Bottou, Carlos Garcia Jurado Suarez, David Grangier, Saleema Amershi, Johan Verwey, and Jina Suh. 2014 · 2014
Earlier work this paper cites.
Proxy tasks and subjective measures can be misleading in evaluating explainable AI systems
Zana Buçinca, Phoebe Lin, Krzysztof Z Gajos, and Elena L Glassman. 2020 · 2020
Earlier work this paper cites.
Design justice: Community-led practices to build the worlds we need
Sasha Costanza-Chock. 2020 · 2020
Earlier work this paper cites.
GPT-3: Its nature, scope, limits, and consequences
Luciano Floridi and Massimo Chiriatti. 2020 · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020 · 2020
Earlier work this paper cites.
Evaluating the Factual Consistency of Abstractive Text Summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for Computational Linguistics, Online, 9332–9346
Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher. 2020 · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Earlier work this paper cites.
Summeval: Re-evaluating summarization evaluation
Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021 · 2021
Earlier work this paper cites.
Manipulating and measuring model interpretability
Forough Poursabzi-Sangdeh, Daniel G Goldstein, Jake M Hofman, Jennifer Wortman Wortman Vaughan, and Hanna Wallach. 2021 · 2021
Earlier work this paper cites.
Scaling Instruction-Finetuned Language Models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022 · 2022
Earlier work this paper cites.
AI Assisted Data Labeling with Interactive Auto Label
Michael Desmond, Michelle Brachman, Evelyn Duesterwald, Casey Dugan, Narendra Nath Joshi, Qian Pan, and Carolina Spina. 2022 · 2022
Earlier work this paper cites.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Cited alongside, same era.
A Closer Look into Using Large Language Models for Automatic Evaluation. In Findings of the Association for Computational Linguistics: EMNLP 2023 , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 8928–8942
Cheng-Han Chiang and Hung-yi Lee. 2023 · 2023
Cited alongside, same era.
Can Large Language Models Be an Alternative to Human Evaluations?. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Association for Computational Linguistics, Toronto, Canada, 15607–15631
Cheng-Han Chiang and Hung yi Lee. 2023 · 2023
Cited alongside, same era.
Retrieval-augmented generation for large language models: A survey
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, and Haofen Wang. 2023 · 2023
Cited alongside, same era.
Alpacafarm: A simulation framework for methods that learn from human feedback
Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S Liang, and Tatsunori B Hashimoto. 2024 · 2024
Closest in time.
HELM: Holistic Evaluation of Language Models
Center for Research on Foundation Models. 2024 · 2024
Closest in time.
Incentives, Framing, and Reliance on Algorithmic Advice: An Experimental Study
Ben Greiner, Philipp Grünwald, Thomas Lindner, Georg Lintner, and Martin Wiernsperger. 2024 · 2024
Closest in time.
An empirical study of llm-as-a-judge for llm evaluation: Fine-tuned judge models are task-specific classifiers
Hui Huang, Yingqi Qu, Jing Liu, Muyun Yang, and Tiejun Zhao. 2024 · 2024
Closest in time.
Llm comparator: Visual analytics for side-by-side evaluation of large language models
Minsuk Kahng, Ian Tenney, Mahima Pushkarna, Michael Xieyang Liu, James Wexler, Emily Reif, Krystal Kallarackal, Minsuk Chang, Michael Terry, and Lucas Dixon. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text
Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam. 2023 · 2023
Cited alongside, same era.
Mistral 7B
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Cited alongside, same era.
Prometheus: Inducing fine-grained evaluation capability in language models
Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, et al · 2023
Cited alongside, same era.
Alpacaeval: An automatic evaluator of instruction-following models
Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023 · 2023
Cited alongside, same era.
G-eval: Nlg evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Cited alongside, same era.
Verbosity bias in preference labeling by large language models
Keita Saito, Akifumi Wachi, Koki Wataoka, and Youhei Akimoto. 2023 · 2023
Cited alongside, same era.
Claude 3.5 sonnet
Anthropic. 2024 · 2024
Cited alongside, same era.
ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis Testing
Ian Arawjo, Chelse Swoopes, Priyan Vaithilingam, Martin Wattenberg, and Elena L Glassman. 2024 · 2024
Cited alongside, same era.
EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24) . Association for Computing Machinery, New York, NY, USA, Article 306, 21 pages
Tae Soo Kim, Yoonjoo Lee, Jamin Shin, Young-Ho Kim, and Juho Kim. 2024a · 2024
Closest in time.
Split and Merge: Aligning Position Biases in LLM-based Evaluators
Zongjie Li, Chaozheng Wang, Pingchuan Ma, Daoyuan Wu, Shuai Wang, Cuiyun Gao, and Yang Liu. 2024 · 2024
Closest in time.
Aligning with human judgement: The role of pairwise preference in large language model evaluators
Yinhong Liu, Han Zhou, Zhijiang Guo, Ehsan Shareghi, Ivan Vulić, Anna Korhonen, and Nigel Collier. 2024 · 2024
Closest in time.
Are self-explanations from Large Language Models faithful?
Andreas Madsen, Sarath Chandar, and Siva Reddy. 2024 · 2024
Closest in time.
Human-Centered Design Recommendations for LLM-as-a-Judge
Qian Pan, Zahra Ashktorab, Michael Desmond, Martin Santillan Cooper, James Johnson, Rahul Nair, Elizabeth Daly, and Werner Geyer. 2024 · 2024
Closest in time.
Constructing Domain-Specific Evaluation Sets for LLM-as-a-judge
Ravi Raju, Swayambhoo Jain, Bo Li, Jonathan Li, and Urmish Thakkar. 2024 · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al · 2024
Closest in time.
Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences
Shreya Shankar, JD Zamfirescu-Pereira, Björn Hartmann, Aditya G Parameswaran, and Ian Arawjo. 2024 · 2024
Closest in time.
Judging the Judges: A Systematic Investigation of Position Bias in Pairwise Comparative Assessments by LLMs
Lin Shi, Chiyu Ma, Wenhua Liang, Weicheng Ma, and Soroush Vosoughi. 2024 · 2024
Closest in time.
Pride and prejudice: LLM amplifies self-bias in self-refinement
Wenda Xu, Guanglei Zhu, Xuandong Zhao, Liangming Pan, Lei Li, and William Wang. 2024 · 2024
Closest in time.
Evaluation of Retrieval-Augmented Generation: A Survey
Hao Yu, Aoran Gan, Kai Zhang, Shiwei Tong, Qi Liu, and Zhaofeng Liu. 2024 · 2024
Closest in time.
Judging LLM-as-a-judge with MT-bench and Chatbot Arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2024 · 2024
Closest in time.
Detecting Hallucinations in Large Language Model Generation: A Token Probability Approach. In Artificial Intelligence and Applications , Hamid R. Arabnia, Leonidas Deligiannidis, Soheyla Amirian, Farzan Shenavarmasouleh, Farid Ghareh Mohammadi, and David de la Fuente (Eds.). Springer Nature Switzerland, Cham, 154–173
Ernesto Quevedo, Jorge Yero Salazar, Rachel Koerner, Pablo Rivas, and Tomas Cerny. 2025 · 2025
Closest in time.