Fetching the paper…
Reading the bibliography…
In this position paper, we argue that human evaluation of generative large language models (LLMs) should be a multidisciplinary undertaking that draws upon insights from disciplines such as user experience research and human behavioral psychology to ensure that the experimental design and results are reliable.
On faithfulness and factuality in abstractive summarization
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020 · 1919
Earlier work this paper cites.
Judgment under uncertainty: Heuristics and biases
Amos Tversky and Daniel Kahneman. 1974 · 1974
Earlier work this paper cites.
The halo effect: Evidence for unconscious alteration of judgments
Richard Nisbett and Timothy Wilson. 1977 · 1977
Earlier work this paper cites.
Sorting-based menu categories
Douglas Hayhoe. 1990 · 1990
Earlier work this paper cites.
Performance vs. preference
Robert W. Bailey. 1993 · 1993
Earlier work this paper cites.
Recommending and evaluating choices in a virtual community of use
Will Hill, Larry Stead, Mark Rosenstein, and George Furnas. 1995 · 1995
Earlier work this paper cites.
Effects of perceptual fluency on judgments of truth
Rolf Reber and Norbert Schwarz. 1999 · 1999
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Evaluating content selection in summarization: The pyramid method
Ani Nenkova and Rebecca Passonneau. 2004 · 2004
Earlier work this paper cites.
A catalog of biases in questionnaires
Bernard C K Choi and Anita W P Pak. 2005 · 2005
Earlier work this paper cites.
on judgments of truth & beauty
Norbert Schwarz. 2006 · 2006
Earlier work this paper cites.
Do People Experience Cognitive Biases while Searching for Information?
Annie Y.S. Lau and Enrico W. Coiera. 2007 · 2007
Earlier work this paper cites.
Comparing rating scales and preference judgements in language evaluation
Anja Belz and Eric Kow. 2010 · 2010
Earlier work this paper cites.
The influence of design aesthetics in usability testing: Effects on user performance and perceived usability
Andreas Sonderegger and Juergen Sauer. 2010 · 2010
Earlier work this paper cites.
A literature review of the anchoring effect
Adrian Furnham and Hua Chu Boo. 2011 · 2011
Earlier work this paper cites.
A comparison of psychometric properties and normality in 4-, 5-, 6-, and 11-point likert scales
Shing-On Leung. 2011 · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
223Processing Fluency, Aesthetic Pleasure, and Culturally Shared Taste
Rolf Reber. 2011 · 2011
Earlier work this paper cites.
When does feeling of fluency matter? how abstract and concrete thinking influence fluency effects
Claire I Tsai and Manoj Thomas. 2011 · 2011
Earlier work this paper cites.
The influence of product aesthetics and usability over the course of time: a longitudinal field experiment
Andreas Uebelbacher Andreas Sonderegger, Gerold Zbinden and Juergen Sauer. 2012 · 2012
Earlier work this paper cites.
The UX Book: Process and guidelines for ensuring a quality user experience
Rex Hartson and Pardha S Pyla. 2012 · 2012
Earlier work this paper cites.
Feeling good and feeling truth: The interactive effects of mood and processing fluency on truth judgments
Alex S. Koch and Joseph P. Forgas. 2012 · 2012
Earlier work this paper cites.
Interrater reliability: the kappa statistic
Mary L McHugh. 2012 · 2012
Earlier work this paper cites.
Designing interfaces for explicit preference elicitation: a user-centered investigation of preference representation and elicitation process
Alina Pommeranz, Joost Broekens, Pascal Wiggers, Willem-Paul Brinkman, and Catholijn M. Jonker. 2012 · 2012
Earlier work this paper cites.
Is beautiful really usable? toward understanding the relation between usability, aesthetics, and affect in hci
Alexandre N. Tuch, Sandra P. Roth, Kasper Hornbæk, Klaus Opwis, and Javier A. Bargas-Avila. 2012 · 2012
Earlier work this paper cites.
Multidimensional quality metrics: a flexible system for assessing translation quality
Aljoscha Burchardt. 2013 · 2013
Earlier work this paper cites.
Power failure: why small sample size undermines the reliability of neuroscience
Katherine S. Button, John P. A. Ioannidis, Claire Mokrysz, Brian A. Nosek, Jonathan Flint, Emma S. J. Robinson, and Marcus R. Munafò. 2013 · 2013
Earlier work this paper cites.
Interrater agreement and interrater reliability: Key concepts, approaches, and applications
Natasa Gisev, J. Simon Bell, and Timothy F. Chen. 2013 · 2013
Earlier work this paper cites.
Chapter 3 - planning
Tom Tullis and Bill Albert. 2013 · 2013
Earlier work this paper cites.
Beliefs and biases in web search
Ryen W. White. 2013 · 2013
Earlier work this paper cites.
Cognitively inspired task design to improve user performance on crowdsourcing platforms
Harini Alagarai Sampath, Rajeev Rajeshuni, and Bipin Indurkhya. 2014 · 2014
Earlier work this paper cites.
The interplay between usability and aesthetics: More evidence for the “what is usable is beautiful” notion
Kai-Christoph Hamborg, Julia Hülsmann, and Kai Kaspar. 2014 · 2014
Earlier work this paper cites.
Big data and large sample size: a cautionary note on the potential for bias
Robert M Kaplan, David A Chambers, and Russell E Glasgow. 2014 · 2014
Earlier work this paper cites.
Research study on importance of usability testing/user experience (ux) testing
M Niranjanamurthy, Archikam Nagaraj, Himaja Gattu, and Puneeth K Shetty. 2014 · 2014
Earlier work this paper cites.
Use and misuse of the likert item responses and other ordinal measures
Phillip A Bishop and Robert L Herron. 2015 · 2015
Earlier work this paper cites.
Is psychology suffering from a replication crisis? what does “failure to replicate” really mean?
Scott E Maxwell, Michael Y Lau, and George S Howard. 2015 · 2015
Earlier work this paper cites.
A formal analysis of the iso 9241-210 definition of user experience
Alexander G. Mirnig, Alexander Meschtscherjakov, Daniela Wurhofer, Thomas Meneweger, and Manfred Tscheligi. 2015 · 2015
Earlier work this paper cites.
1,500 scientists lift the lid on reproducibility
Monya Baker. 2016 · 2016
Cited alongside, same era.
Another look at likert scales
Fern K Willits, Gene L Theodori, and AE Luloff. 2016 · 2016
Cited alongside, same era.
Inter-annotator Agreement , pages 297–313. Springer Netherlands, Dordrecht
Ron Artstein. 2017 · 2017
Cited alongside, same era.
Clarity is a worthwhile quality: On the role of task clarity in microtask crowdsourcing
Ujwal Gadiraju, Jie Yang, and Alessandro Bozzon. 2017 · 2017
Cited alongside, same era.
The interplay of cognition and feelings: Fluency
Rainer Greifeneder and Herbert Bless. 2017 · 2017
Cited alongside, same era.
An overview of interrater agreement on likert scales for researchers and practitioners
Thomas A. O’Neill. 2017 · 2017
Cited alongside, same era.
Quality control questions on amazon’s mechanical turk (mturk): A randomized trial of impact on the usaudit, phq-9, and gad-7
Jon Agley, Yunyu Xiao, Rachael Nolan, and Lilian Golzarri-Arroyo. 2022 · 2022
Later among the works it cites.
Re-examining system-level correlations of automatic summarization evaluation metrics
Daniel Deutsch, Rotem Dror, and Dan Roth. 2022 · 2022
Later among the works it cites.
Accelerating human authorship of information extraction rules
Dayne Freitag, John Cadigan, John Niekrasz, and Robert Sasseen. 2022 · 2022
Later among the works it cites.
DialSummEval: Revisiting summarization evaluation for dialogues
Mingqi Gao and Xiaojun Wan. 2022 · 2022
Later among the works it cites.
Improved recommender systems by denoising ratings in highly sparse datasets through individual rating confidence
Nima Joorabloo, Mahdi Jalili, and Yongli Ren. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Confusing the crowd: Task instruction quality on amazon mechanical turk
Meng-Han Wu and Alexander Quinn. 2017 · 2017
Cited alongside, same era.
Cognitive biases in crowdsourcing
Carsten Eickhoff. 2018 · 2018
Cited alongside, same era.
So, What Are Cognitive Biases? , pages 1–10. Springer International Publishing, Cham
Geoffrey Ellis. 2018 · 2018
Cited alongside, same era.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018 · 2018
Cited alongside, same era.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018 · 2018
Cited alongside, same era.
Towards designing unbiased replication studies in information visualization
Poorna Talkad Sukumar and Ronald Metoyer. 2018 · 2018
Cited alongside, same era.
Rating consistency is consistently underrated: an exploratory analysis of movie-tag rating inconsistency
Denis Kotkov, Alan Medlar, Umesh Raj Satyal, Alexandr V. Maslov, Mats Neovius, and Dorota Glowacka. 2022 · 2022
Later among the works it cites.
Opportunities for human-centered evaluation of machine translation systems
Daniel Liebling, Katherine Heller, Samantha Robertson, and Wesley Deng. 2022 · 2022
Later among the works it cites.
Data contamination: From memorization to exploitation
Inbal Magar and Roy Schwartz. 2022 · 2022
Later among the works it cites.
Red teaming language models with language models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. 2022 · 2022
Later among the works it cites.
In search of ambiguity: A three-stage workflow design to clarify annotation guidelines for crowd workers
Vivek Krishna Pradhan, Mike Schaekermann, and Matthew Lease. 2022 · 2022
Later among the works it cites.
Two contrasting data annotation paradigms for subjective NLP tasks
Paul Rottger, Bertie Vidgen, Dirk Hovy, and Janet Pierrehumbert. 2022 · 2022
Later among the works it cites.
On the robustness of offensive language classifiers
Jonathan Rusert, Zubair Shafiq, and Padmini Srinivasan. 2022 · 2022
Later among the works it cites.
Non-repeatable experiments and non-reproducible results: The reproducibility crisis in human evaluation in NLP
Anya Belz, Craig Thomson, Ehud Reiter, and Simon Mille. 2023 · 2023
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Big bench authors. 2023 · 2023
Later among the works it cites.
A customized text sanitization mechanism with differential privacy
Sai Chen, Fengran Mo, Yanhao Wang, Cen Chen, Jian-Yun Nie, Chengyu Wang, and Jamie Cui. 2023 · 2023
Later among the works it cites.
A closer look into using large language models for automatic evaluation
Cheng-Han Chiang and Hung-yi Lee. 2023 · 2023
Later among the works it cites.
Principles from clinical research for nlp model generalization
Aparna Elangovan, Jiayuan He, Yuan Li, and Karin Verspoor. 2023 · 2023
Later among the works it cites.
ROBBIE: Robust bias evaluation of large generative language models
David Esiobu, Xiaoqing Tan, Saghar Hosseini, Megan Ung, Yuchen Zhang, Jude Fernandes, Jane Dwivedi-Yu, Eleonora Presani, Adina Williams, and Eric Smith. 2023 · 2023
Later among the works it cites.
Lawbench: Benchmarking legal knowledge of large language models
Zhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou, Zhuo Han, Songyang Zhang, Kai Chen, Zongwen Shen, and Jidong Ge. 2023 · 2023
Later among the works it cites.
Reduce human labor on evaluating conversational information retrieval system: A human-machine collaboration approach
Chen Huang, Peixin Qin, Wenqiang Lei, and Jiancheng Lv. 2023 · 2023
Later among the works it cites.
NamHyeok Kim and Chanjun Park. 2023 · 2023
Later among the works it cites.
Chatgpt: Jack of all trades, master of none
Jan Kocoń, Igor Cichecki, Oliwier Kaszyca, Mateusz Kochanek, Dominika Szydło, Joanna Baran, Julita Bielaniewicz, Marcin Gruza, Arkadiusz Janz, Kamil Kanclerz, Anna Kocoń, Bartłomiej Koptyra, Wiktoria Mieleszczenko-Kowszewicz, Piotr Miłkowski, Marcin Oleksy, Maciej Piasecki, Łukasz Radliński, Konrad Wojtasik, Stanisław Woźniak, and Przemysław Kazienko. 2023 · 2023
Later among the works it cites.
Multi-step jailbreaking privacy attacks on ChatGPT
Haoran Li, Dadi Guo, Wei Fan, Mingshi Xu, Jie Huang, Fanpu Meng, and Yangqiu Song. 2023 · 2023
Later among the works it cites.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Alexander Cosgrove, Christopher D Manning, Christopher Re, Diana Acosta-Navas, Drew Arad Hudson, Eric Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue WANG, Keshav Santhanam, Laurel Orr, Lucia Zheng, Mert Yuksekgonul, Mirac Suzgun, Nathan Kim, Neel Guha, Niladri S. Chatterji, Omar Khattab, Peter Henderson, Qian Huang, Ryan Andrew Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas Icard, Tianyi Zhang, Vishrav Chaudhary, William Wang, Xuechen Li, Yifan Mai, Yuhui Zhang, and Yuta Koreeda. 2023 · 2023
Later among the works it cites.
LLM-eval: Unified multi-dimensional automatic evaluation for open-domain conversations with large language models
Yen-Ting Lin and Yun-Nung Chen. 2023 · 2023
Later among the works it cites.
Revisiting the gold standard: Grounding summarization evaluation with robust human evaluation
Yixin Liu, Alex Fabbri, Pengfei Liu, Yilun Zhao, Linyong Nan, Ruilin Han, Simeng Han, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, and Dragomir Radev. 2023 · 2023
Later among the works it cites.
ASSERT: Automated safety scenario red teaming for evaluating the robustness of large language models
Alex Mei, Sharon Levy, and William Wang. 2023 · 2023
Later among the works it cites.
Statistically controlling for processing fluency reduces the aesthetic-usability effect
Jan Preßler, Lukas Schmid, and Jörn Hurtienne. 2023 · 2023
Later among the works it cites.
AART: AI-assisted red-teaming with diverse data generation for new LLM-powered applications
Bhaktipriya Radharapu, Kevin Robinson, Lora Aroyo, and Preethi Lahoti. 2023 · 2023
Later among the works it cites.
NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark
Oscar Sainz, Jon Campos, Iker García-Ferrero, Julen Etxaniz, Oier Lopez de Lacalle, and Eneko Agirre. 2023 · 2023
Later among the works it cites.
What’s the meaning of superhuman performance in today’s NLU?
Simone Tedeschi, Johan Bos, Thierry Declerck, Jan Hajič, Daniel Hershcovich, Eduard Hovy, Alexander Koller, Simon Krek, Steven Schockaert, Rico Sennrich, Ekaterina Shutova, and Roberto Navigli. 2023 · 2023
Later among the works it cites.
Using membership inference attacks to evaluate privacy-preserving language modeling fails for pseudonymizing data
Thomas Vakili and Hercules Dalianis. 2023 · 2023
Later among the works it cites.
K-alpha calculator–krippendorff’s alpha calculator: A user-friendly tool for computing krippendorff’s alpha inter-rater reliability coefficient
Giacomo Marzi, Marco Balzano, and Davide Marchiori. 2024 · 2024
Closest in time.
Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting
Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. 2024 · 2024
Closest in time.
Trustllm: Trustworthiness in large language models
Lichao Sun, Yue Huang, Haoran Wang, Siyuan Wu, Qihui Zhang, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, Zhengliang Liu, Yixin Liu, Yijue Wang, Zhikun Zhang, Bhavya Kailkhura, Caiming Xiong, Chaowei Xiao, Chunyuan Li, Eric Xing, Furong Huang, Hao Liu, Heng Ji, Hongyi Wang, Huan Zhang, Huaxiu Yao, Manolis Kellis, Marinka Zitnik, Meng Jiang, Mohit Bansal, James Zou, Jian Pei, Jian Liu, Jianfeng Gao, Jiawei Han, Jieyu Zhao, Jiliang Tang, Jindong Wang, John Mitchell, Kai Shu, Kaidi Xu, Kai-Wei Chang, Lifang He, Lifu Huang, Michael Backes, Neil Zhenqiang Gong, Philip S. Yu, Pin-Yu Chen, Quanquan Gu, Ran Xu, Rex Ying, Shuiwang Ji, Suman Jana, Tianlong Chen, Tianming Liu, Tianyi Zhou, William Wang, Xiang Li, Xiangliang Zhang, Xiao Wang, Xing Xie, Xun Chen, Xuyu Wang, Yan Liu, Yanfang Ye, Yinzhi Cao, Yong Chen, and Yue Zhao. 2024 · 2024
Closest in time.
https://www.vocabulary.com/dictionary/differentiate
vocabulary.com · 2024
Closest in time.