Fetching the paper…
Reading the bibliography…
Across academia, industry, and government, there is an increasing awareness that the measurement tasks involved in evaluating generative AI (GenAI) systems are especially difficult.
Content analysis in communication research, 1952
Bernard Berelson · 1952
Earlier work this paper cites.
Construct validity in psychological tests
Lee J Cronbach and Paul E Meehl · 1955
Earlier work this paper cites.
The nature and origins of mass opinion
John Zaller · 1992
Earlier work this paper cites.
Validity and washback in language testing
Samuel Messick · 1996
Earlier work this paper cites.
Measurement validity: A shared standard for qualitative and quantitative research
Robert Adcock and David Collier · 2001
Earlier work this paper cites.
Privacy is an essentially contested concept: a multi-dimensional analytic for mapping privacy
Deirdre K Mulligan, Colin Koopman, and Nick Doty · 2016
Earlier work this paper cites.
Hate speech in public discourse: A pessimistic defense of counterspeech
Maxime Lepoutre · 2017
Earlier work this paper cites.
This thing called fairness: Disciplinary confusion realizing a value in technology
Deirdre K. Mulligan, Joshua A. Kroll, Nitin Kohli, and Richmond Y. Wong · 2019
Earlier work this paper cites.
Language (technology) is power: A critical survey of ‘bias’ in nlp
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach · 2020
Cited alongside, same era.
Stereoset: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy · 2020
Cited alongside, same era.
Crows-pairs: A challenge dataset for measuring social biases in masked language models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R Bowman · 2020
Cited alongside, same era.
Sociolinguistically Driven Approaches for Just Natural Language Processing
Su Lin Blodgett · 2021
Cited alongside, same era.
Stereotyping norwegian salmon: An inventory of pitfalls in fairness benchmark datasets
Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach · 2021
Cited alongside, same era.
Making Intelligence: Ethical Values in IQ and ML Benchmarks
Borhane Blili-Hamelin and Leif Hancox-Li · 2023
Later among the works it cites.
DALL-Eval: Probing the Reasoning Skills and Social Biases of Text-to-Image Generation Models, 2023
Jaemin Cho, Abhay Zala, and Mohit Bansal · 2023
Later among the works it cites.
Report of the 1st Workshop on Generative AI and Law
A. Feder Cooper, Katherine Lee, James Grimmelmann, Daphne Ippolito, Christopher Callison-Burch, Christopher A. Choquette-Choo, Niloofar Mireshghallah, Miles Brundage, David Mimno, Madiha Zahrah Choksi, Jack M. Balkin, Nicholas Carlini, Christopher De Sa, Jonathan Frankle, Deep Ganguli, Bryant Gipson, Andres Guadamuz, Swee Leng Harris, Abigail Z. Jacobs, Elizabeth Joh, Gautam Kamath, Mark Lemley, Cass Matthews, Christine McLeavey, Corynne McSherry, Milad Nasr, Paul Ohm, Adam Roberts, Tom Rubin, Pamela Samuelson, Ludwig Schubert, Kristen Vaccaro, Luis Villa, Felix Wu, and Elana Zeide · 2023
Later among the works it cites.
Evaluating genera-purpose AI with psychometrics
Xiting Wang, Liming Jiang, Jose Hernandez-Orallo, David Stillwell, Luning Sun, Fang Luo, and Xing Xie · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Emergent Unfairness in Algorithmic Fairness-Accuracy Trade-Off Research
A. Feder Cooper, Ellen Abrams, and NA NA · 2021
Cited alongside, same era.
Measurement and fairness
Abigail Z Jacobs and Hanna Wallach · 2021
Cited alongside, same era.
Red Teaming Language Models with Language Models, 2022
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving · 2022
Cited alongside, same era.
YouTube Hate Speech Policy
URL https://support.google.com/youtube/answer/2801939?hl=en
Cited in the paper.
Later among the works it cites.
Representational harms through the lens of speech act theory
Emily Corvi, Hannah Washington, Stefanie Reed, Chad Atalla, Alexandra Chouldechova, Alex Dow, Jean Garcia-Gathright, Nicholas Pangakis, Emily Sheng, Dan Vann, Matthew Vogel, and Hanna Wallach · 2024
Closest in time.
Visage: A global-scale analysis of visual stereotypes in text-to-image generation, 2024
Akshita Jha, Vinodkumar Prabhakaran, Remi Denton, Sarah Laszlo, Shachi Dave, Rida Qadri, Chandan K. Reddy, and Sunipa Dev · 2024
Closest in time.
ECBD: Evidence-centered benchmark design for NLP
Yu Lu Liu, Su Lin Blodgett, Jackie Cheung, Q. Vera Liao, Alexandra Olteanu, and Ziang Xiao · 2024
Closest in time.
Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, 2024
National Institute for Standards and Technology · 2024
Closest in time.