Fetching the paper…
Reading the bibliography…
The measurement tasks involved in evaluating generative AI (GenAI) systems lack sufficient scientific rigor, leading to what has been described as "a tangle of sloppy tests [and] apples-to-oranges comparisons" (Roose, 2024).
Content Analysis in Communication Research
Berelson, B · 1953
Earlier work this paper cites.
Construct validity in psychological tests
Cronbach, L. J. and Meehl, P. E · 1955
Earlier work this paper cites.
Architecture of the IBM System/360
Amdahl, G. M., Blaauw, G. A., and Brooks, F. P · 1964
Earlier work this paper cites.
Algorithm = Logic + Control
Kowalski, R · 1979
Earlier work this paper cites.
Validity
Messick, S · 1987
Earlier work this paper cites.
The nature and origins of mass opinion
Zaller, J · 1992
Earlier work this paper cites.
Dictionary of Psychology
Cardwell, M · 1996
Earlier work this paper cites.
Validity and washback in language testing
Messick, S · 1996
Earlier work this paper cites.
Measurement validity: A shared standard for qualitative and quantitative research
Adcock, R. and Collier, D · 2001
Earlier work this paper cites.
Meaurement Theory and Practice
Hand, D. J · 2004
Earlier work this paper cites.
Coder reliability and misclassification in the human coding of party manifestos
Mikhaylov, S., Laver, M., and Benoit, K. R · 2012
Earlier work this paper cites.
Privacy is an essentially contested concept: A multi-dimensional analytic for mapping privacy
Mulligan, D. K., Koopman, C., and Doty, N · 2016
Earlier work this paper cites.
This thing called fairness: Disciplinary confusion realizing a value in technology
Mulligan, D. K., Kroll, J. A., Kohli, N., and Wong, R. Y · 2019
Earlier work this paper cites.
Roles for computing in social change
Abebe, R., Barocas, S., Kleinberg, J., Levy, K., Raghavan, M., and Robinson, D. G · 2020
Earlier work this paper cites.
Language (technology) is power: A critical survey of “bias” in NLP
Blodgett, S. L., Barocas, S., Daumé III, H., and Wallach, H · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
CrowS-Pairs: A challenge dataset for measuring social biases in masked language models
Nangia, N., Vania, C., Bhalerao, R., and Bowman, S. R · 2020
Earlier work this paper cites.
Stereotyping Norwegian salmon: An inventory of pitfalls in fairness benchmark datasets
Blodgett, S. L., Lopez, G., Olteanu, A., Sim, R., and Wallach, H · 2021
Earlier work this paper cites.
Extracting training data from large language models
Carlini, N., Tramèr, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, Ú., Oprea, A., and Raffel, C · 2021
Earlier work this paper cites.
Emergent unfairness in algorithmic fairness-accuracy trade-off research
Cooper, A. F., Abrams, E., and NA, N · 2021
Earlier work this paper cites.
Intrinsic bias metrics do not correlate with application bias
Goldfarb-Tarrant, S., Marchant, R., Muñoz Sánchez, R., Pandya, M., and Lopez, A · 2021
Earlier work this paper cites.
Measurement as governance in and for responsible AI
Jacobs, A. Z · 2021
Earlier work this paper cites.
Measurement and fairness
Jacobs, A. Z. and Wallach, H · 2021
Earlier work this paper cites.
StereoSet: Measuring stereotypical bias in pretrained language models
Nadeem, M., Bethke, A., and Reddy, S · 2021
Cited alongside, same era.
AI and the everything in the whole wide world benchmark
Raji, D., Denton, E., Bender, E. M., Hanna, A., and Paullada, A · 2021
Cited alongside, same era.
Gender bias in machine translation
Savoldi, B., Gaido, M., Bentivogli, L., Negri, M., and Turchi, M · 2021
Cited alongside, same era.
Evaluation gaps in machine learning practice
Hutchinson, B., Rostamzadeh, N., Greer, C., Heller, K., and Prabhakaran, V · 2022
Cited alongside, same era.
Deduplicating training data makes language models better
Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C., and Carlini, N · 2022
Cited alongside, same era.
Red teaming language models with language models
Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G · 2022
ECBD: Evidence-centered benchmark design for NLP
Liu, Y. L., Blodgett, S. L., Cheung, J., Liao, Q. V., Olteanu, A., and Xiao, Z · 2024
Later among the works it cites.
The AI Index 2024 Annual Report
Maslej, N., Fattorini, L., Perrault, R., Parli, V., Reuel, A., Brynjolfsson, E., Etchemendy, J., Ligett, K., Lyons, T., Manyika, J., Niebles, J. C., Shoham, Y., Wald, R., and Clark, J · 2024
Later among the works it cites.
Artificial intelligence Risk Management Framework: Generative Artificial Intelligence Profile, 2024
National Institute for Standards and Technology · 2024
Later among the works it cites.
Gaps in the safety evaluation of generative AI
Rauh, M., Marchal, N., Manzini, A., Hendricks, L. A., Comanescu, R., Akbulut, C., Stepleton, T., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., Gabriel, I., Rieser, V., Isaac, W., and Weidinger, L · 2024
Later among the works it cites.
A.I. has a measurement problem
Roose, K · 2024
Later among the works it cites.
Undesirable biases in NLP: Addressing challenges of measurement
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Measuring representational harms in image captioning
Wang, A., Barocas, S., Laird, K., and Wallach, H · 2022
Cited alongside, same era.
Making intelligence: Ethical values in IQ and ML benchmarks
Blili-Hamelin, B. and Hancox-Li, L · 2023
Cited alongside, same era.
Report of the 1st Workshop on Generative AI and Law
Cooper, A. F., Lee, K., Grimmelmann, J., Ippolito, D., Callison-Burch, C., Choquette-Choo, C. A., Mireshghallah, N., Brundage, M., Mimno, D., Choksi, M. Z., Balkin, J. M., Carlini, N., Sa, C. D., Frankle, J., Ganguli, D., Gipson, B., Guadamuz, A., Harris, S. L., Jacobs, A. Z., Joh, E., Kamath, G., Lemley, M., Matthews, C., McLeavey, C., McSherry, C., Nasr, M., Ohm, P., Roberts, A., Rubin, T., Samuelson, P., Schubert, L., Vaccaro, K., Villa, L., Wu, F., and Zeide, E · 2023
Cited alongside, same era.
Taxonomizing and measuring representational harms: A look at image tagging
Katzman, J., Wang, A., Scheuerman, M., Blodgett, S. L., Laird, K., Wallach, H., and Barocas, S · 2023
Cited alongside, same era.
Holistic evaluation of language models
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., Newman, B., Yuan, B., Yan, B., Zhang, C., Cosgrove, C. A., Manning, C. D., Re, C., Acosta-Navas, D., Hudson, D. A., Zelikman, E., Durmus, E., Ladhak, F., Rong, F., Ren, H., Yao, H., WANG, J., Santhanam, K., Orr, L., Zheng, L., Yuksekgonul, M., Suzgun, M., Kim, N., Guha, N., Chatterji, N. S., Khattab, O., Henderson, P., Huang, Q., Chi, R. A., Xie, S. M., Santurkar, S., Ganguli, S., Hashimoto, T., Icard, T., Zhang, T., Chaudhary, V., Wang, W., Li, X., Mai, Y., Zhang, Y., and Koreeda, Y · 2023
Cited alongside, same era.
Diffusion art or digital forgery? Investigating data replication in diffusion models
Somepalli, G., Singla, V., Goldblum, M., Geiping, J., and Goldstein, T · 2023
Cited alongside, same era.
van der Wal, O., Bachmann, D., Leidinger, A., van Maanen, L., Zuidema, W., and Schulz, K · 2024
Later among the works it cites.
Position: Measure dataset diversity, don’t just claim it
Zhao, D., Andrews, J., Papakyriakopoulos, O., and Xiang, A · 2024
Later among the works it cites.
Measuring non-adversarial reproduction of training data in large language models
Aerni, M., Rando, J., Debenedetti, E., Carlini, N., Ippolito, D., and Tramèr, F · 2025
Closest in time.
Position: Medical large language model benchmarks should prioritize construct validity
Alaa, A., Hartvigsen, T., Golchini, N., Dutta, S., Dean, F., Raji, I. D., and Zack, T · 2025
Closest in time.
How to build a better AI benchmark
Brandom, R · 2025
Closest in time.
The files are in the computer: Copyright, memorization, and generative AI
Cooper, A. F. and Grimmelmann, J · 2025
Closest in time.
Taxonomizing representational harms using speech act theory
Corvi, E., Washington, H., Reed, S., Atalla, C., Chouldechova, A., Dow, P. A., Garcia-Gathright, J., Pangakis, N., Sheng, E., Vann, D., et al · 2025
Closest in time.
Can we trust AI benchmarks? An interdisciplinary review of current issues in AI evaluation
Eriksson, M., Purificato, E., Noroozian, A., Vinagre, J., Chaslot, G., Gomez, E., and Fernandez-Llorca, D · 2025
Closest in time.
Measuring memorization in language models via probabilistic extraction
Hayes, J., Swanberg, M., Chaudhari, H., Yona, I., Shumailov, I., Nasr, M., Choquette-Choo, C. A., Lee, K., and Cooper, A. F · 2025
Closest in time.
Talkin’ ’bout AI generation: Copyright and the generative-AI supply chain
Lee, K., Cooper, A. F., and Grimmelmann, J · 2025
Closest in time.
Position: Rethinking LLM bias probing using lessons from the social sciences
Morehouse, K. N., Swaroop, S., and Pan, W · 2025
Closest in time.
Scalable extraction of training data from aligned, production language models
Nasr, M., Rando, J., Carlini, N., Hayase, J., Jagielski, M., Cooper, A. F., Ippolito, D., Choquette-Choo, C. A., Tramèr, F., and Lee, K · 2025
Closest in time.
Machine learners should acknowledge the legal implications of large language models as personal data
Nolte, H., Finck, M., and Meding, K · 2025
Closest in time.
Keeping humans in the loop: Human-centered automated annotation with generative AI
Pangakis, N. and Wolken, S · 2025
Closest in time.
Recite, reconstruct, recollect: Memorization in LMs as a multifaceted phenomenon
Prashanth, U. S., Deng, A., O’Brien, K., V, J. S., Khan, M. A., Borkar, J., Choquette-Choo, C. A., Fuehne, J. R., Biderman, S., Ke, T., Lee, K., and Saphra, N · 2025
Closest in time.
Measurement to meaning: A validity-centered framework for AI evaluation
Salaudeen, O., Reuel, A., Ahmed, A., Bedi, S., Robertson, Z., Sundar, S., Domingue, B., Wang, A., and Koyejo, S · 2025
Closest in time.
End-to-end arguments in system design
Saltzer, J. H., Reed, D. P., and Clark, D. D · 2071
Closest in time.