Fetching the paper…
Reading the bibliography…
The recent development of generative large language models (LLMs) poses new challenges for model evaluation that the research community and industry have been grappling with.
The psychology of human-computer interaction. usa: L, 1983
Card, S. K., Newell, A., and Moran, T. P · 1983
Earlier work this paper cites.
ACM SIGCHI curricula for human-computer interaction
Hewett, T. T., Baecker, R., Card, S., Carey, T., Gasen, J., Mantei, M., Perlman, G., Strong, G., and Verplank, W · 1992
Earlier work this paper cites.
Methodology matters: Doing research in the behavioral and social sciences
McGrath, J. E · 1995
Earlier work this paper cites.
How to conduct a heuristic evaluation
Nielsen, J · 1995
Earlier work this paper cites.
The growth of cognitive modeling in human-computer interaction since goms
Olson, J. R. and Olson, G. M · 1995
Earlier work this paper cites.
Cognitive walkthroughs
Lewis, C. and Wharton, C · 1997
Earlier work this paper cites.
The intellectual challenge of cscw: The gap between social requirements and technical feasibility
Ackerman, M. S · 2000
Earlier work this paper cites.
What is ecological validity? a dimensional analysis
Schmuckler, M. A · 2001
Earlier work this paper cites.
Designing interaction, not interfaces
Beaudouin-Lafon, M · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
From mice to men-24 years of evaluation in chi
Barkhuus, L. and Rode, J. A · 2007
Earlier work this paper cites.
Evaluating user interface systems research
Olsen Jr, D. R · 2007
Earlier work this paper cites.
Cognitive models of human-information interaction
Pirolli, P · 2007
Earlier work this paper cites.
Usability evaluation considered harmful (some of the time)
Greenberg, S. and Buxton, B · 2008
Earlier work this paper cites.
Stories from the field: Reflections on hci4d experiences
Anokwa, Y., Smyth, T. N., Ramachandran, D., Sherwani, J., Schwartzman, Y., Luk, R., Ho, M., Moraveji, N., and DeRenzi, B · 2009
Earlier work this paper cites.
An automatic dialog simulation technique to develop and evaluate interactive conversational agents
Griol, D., Carbó, J., and Molina, J. M · 2013
Earlier work this paper cites.
Changing perspectives on evaluation in hci: past, present, and future
MacDonald, C. M. and Atwood, M. E · 2013
Earlier work this paper cites.
Ways of Knowing in HCI , volume 2
Olson, J. S. and Kellogg, W. A · 2014
Earlier work this paper cites.
Teaching machines to read and comprehend
Hermann, K. M., Kociský, T., Grefenstette, E., Espeholt, L., Kay, W., Suleyman, M., and Blunsom, P · 2015
Earlier work this paper cites.
Abstractive text summarization using sequence-to-sequence rnns and beyond
Nallapati, R., Zhou, B., Gulcehre, C., Xiang, B., et al · 2016
Cited alongside, same era.
Towards a rigorous science of interpretable machine learning
Doshi-Velez, F. and Kim, B · 2017
Cited alongside, same era.
Samsum corpus: A human-annotated dialogue dataset for abstractive summarization
Gliwa, B., Mochol, I., Biesek, M., and Wawer, A · 2019
Cited alongside, same era.
Ask not what ai can do, but what ai should do: Towards a framework of task delegability
Lubars, B. and Tan, C · 2019
Cited alongside, same era.
Fairness and abstraction in sociotechnical systems
Selbst, A. D., Boyd, D., Friedler, S. A., Venkatasubramanian, S., and Vertesi, J · 2019
Cited alongside, same era.
Connecting algorithmic research and usage contexts: A perspective of contextualized evaluation for explainable ai
Liao, Q. V., Zhang, Y., Luss, R., Doshi-Velez, F., and Dhurandhar, A · 2022
Later among the works it cites.
Dropping the gre, keeping the gre, or gre-optional admissions? considering tradeoffs and fairness
Newman, D. A., Tang, C., Song, Q. C., and Wee, S · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
A survey of evaluation metrics used for nlg systems
Sai, A. B., Mohankumar, A. K., and Khapra, M. M · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bertscore: Evaluating text generation with bert
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y · 2019
Cited alongside, same era.
Human factors in model interpretability: Industry practices, challenges, and needs
Hong, S. R., Hullman, J., and Bertini, E · 2020
Cited alongside, same era.
Twenty years of confusion in human evaluation: Nlg needs evaluation sheets and standardised definition
Howcroft, D., Belz, A., Clinciu, M., Gkatzia, D., Hasan, S. A., Mahamood, S., Mille, S., Van Miltenburg, E., Santhanam, S., and Rieser, V · 2020
Cited alongside, same era.
Questioning the ai: informing design practices for explainable ai user experiences
Liao, Q. V., Gruen, D., and Miller, S · 2020
Cited alongside, same era.
A human-centered agenda for intelligible machine learning
Vaughan, J. W. and Wallach, H · 2020
Cited alongside, same era.
Evaluating conversational recommender systems via user simulation
Zhang, S. and Balog, K · 2020
Cited alongside, same era.
The reprogen shared task on reproducibility of human evaluations in nlg: Overview and results
Belz, A., Shimorina, A., Agarwal, S., and Reiter, E · 2021
Cited alongside, same era.
Later among the works it cites.
Taxonomy of risks posed by language models
Weidinger, L., Uesato, J., Rauh, M., Griffin, C., Huang, P.-S., Mellor, J., Glaese, A., Cheng, M., Balle, B., Kasirzadeh, A., et al · 2022
Later among the works it cites.
Deconstructing nlg evaluation: Evaluation practices, assumptions, and their implications
Zhou, K., Blodgett, S. L., Trischler, A., Daumé III, H., Suleman, K., and Olteanu, A · 2022
Later among the works it cites.
Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text
Gehrmann, S., Clark, E., and Sellam, T · 2023
Closest in time.
Evaluating large language models in generating synthetic hci research data: a case study
Hämäläinen, P., Tavast, M., and Kunnari, A · 2023
Closest in time.
Towards a science of human-ai decision making: An overview of design space in empirical human-subject studies
Lai, V., Chen, C., Smith-Renner, A., Liao, Q. V., and Tan, C · 2023
Closest in time.
Holistic evaluation of language models
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., Newman, B., Yuan, B., Yan, B., Zhang, C., Cosgrove, C. A., Manning, C. D., Re, C., Acosta-Navas, D., Hudson, D. A., Zelikman, E., Durmus, E., Ladhak, F., Rong, F., Ren, H., Yao, H., WANG, J., Santhanam, K., Orr, L., Zheng, L., Yuksekgonul, M., Suzgun, M., Kim, N., Guha, N., Chatterji, N. S., Khattab, O., Henderson, P., Huang, Q., Chi, R. A., Xie, S. M., Santurkar, S., Ganguli, S., Hashimoto, T., Icard, T., Zhang, T., Chaudhary, V., Wang, W., Li, X., Mai, Y., Zhang, Y., and Koreeda, Y · 2023
Closest in time.
Designerly understanding: Information needs for model transparency to support design ideation for ai-powered user experience
Liao, Q. V., Subramonyam, H., Wang, J., and Wortman Vaughan, J · 2023
Closest in time.
Humans and algorithms work together—so study them together
Matias, J. N · 2023
Closest in time.
Evaluating evaluation metrics: A framework for analyzing nlg evaluation metrics using measurement theory
Xiao, Z., Zhang, S., Lai, V., and Liao, Q. V · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2023
Closest in time.
The illusion of artificial inclusion
Agnew, W., Bergman, A. S., Chien, J., Díaz, M., El-Sayed, S., Pittman, J., Mohamed, S., and McKee, K. R · 2024
Closest in time.
ECBD: Evidence-centered benchmark design for NLP
Liu, Y. L., Blodgett, S. L., Cheung, J., Liao, Q. V., Olteanu, A., and Xiao, Z · 2024
Closest in time.
Dissociating language and thought in large language models
Mahowald, K., Ivanova, A. A., Blank, I. A., Kanwisher, N., Tenenbaum, J. B., and Fedorenko, E · 2024
Closest in time.
Analysing utterances in llm-based user simulation for conversational search
Sekulić, I., Alinannejadi, M., and Crestani, F · 2024
Closest in time.