Fetching the paper…
Reading the bibliography…
The valid measurement of generative AI (GenAI) systems' capabilities, risks, and impacts forms the bedrock of our ability to evaluate these systems.
Measurement validity: A shared standard for qualitative and quantitative research
Robert Adcock and David Collier · 2001
Earlier work this paper cites.
Show Your Work: Improved Reporting of Experimental Results
Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, and Noah A. Smith · 2019
Earlier work this paper cites.
Dissecting racial bias in an algorithm used to manage the health of populations
Ziad Obermeyer, Brian Powers, Christine Vogeli, and Sendhil Mullainathan · 2019
Earlier work this paper cites.
Problem formulation and fairness
Samir Passi and Solon Barocas · 2019
Earlier work this paper cites.
Counterfactual risk assessments, evaluation, and fairness
Amanda Coston, Alan Mishler, Edward H Kennedy, and Alexandra Chouldechova · 2020
Earlier work this paper cites.
Stereotyping norwegian salmon: An inventory of pitfalls in fairness benchmark datasets
Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach · 2021
Earlier work this paper cites.
Hyperparameter Optimization Is Deceiving Us, and How to Stop It
A. Feder Cooper, Yucheng Lu, Jessica Forde, and Christopher M De Sa · 2021
Earlier work this paper cites.
Proxies: The cultural work of standing in
Dylan Mulvin · 2021
Cited alongside, same era.
Are my deep learning systems fair? an empirical study of fixed-seed training
Shangshu Qian, Hung Viet Pham, Thibaud Lutellier, Zeou Hu, Jungwon Kim, Lin Tan, Yaoliang Yu, Jiahao Chen, and Sameena Shah · 2021
Cited alongside, same era.
Quantifying memorization across neural language models
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang · 2022
Cited alongside, same era.
Preventing Verbatim Memorization in Language Models Gives a False Sense of Privacy, 2023
Daphne Ippolito, Florian Tramèr, Milad Nasr, Chiyuan Zhang, Matthew Jagielski, Katherine Lee, Christopher A. Choquette-Choo, and Nicholas Carlini · 2023
Cited alongside, same era.
Optimization’s neglected normative commitments
Benjamin Laufer, Thomas Gilbert, and Helen Nissenbaum · 2023
Cited alongside, same era.
The Files are in the Computer: Copyright, Memorization, and Generative AI
A Feder Cooper and James Grimmelmann · 2024
Closest in time.
Arbitrariness and Social Prediction: The Confounding Role of Variance in Fair Classification
A. Feder Cooper, Katherine Lee, Madiha Zahrah Choksi, Solon Barocas, Christopher De Sa, James Grimmelmann, Jon Kleinberg, Siddhartha Sen, and Baobao Zhang · 2024
Closest in time.
ECBD: Evidence-centered benchmark design for NLP
Yu Lu Liu, Su Lin Blodgett, Jackie Cheung, Q. Vera Liao, Alexandra Olteanu, and Ziang Xiao · 2024
Closest in time.
The AI index 2024 annual report
Nestor Maslej, Loredana Fattorini, Raymond Perrault, Vanessa Parli, Anka Reuel, Erik Brynjolfsson, John Etchemendy, Katrina Ligett, Terah Lyons, James Manyika, Juan Carlos Niebles, Yoav Shoham, Russell Wald, and Jack Clark · 2024
Closest in time.
On the consistency of hyper-parameter selection in value-based deep reinforcement learning, 2024
Johan Obando-Ceron, João G. M. Araújo, Aaron Courville, and Pablo Samuel Castro · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Evaluating the social impact of generative ai systems in systems and society
Irene Solaiman, Zeerak Talat, William Agnew, Lama Ahmad, Dylan Baker, Su Lin Blodgett, Canyu Chen, Hal Daumé III, Jesse Dodge, Isabella Duan, et al · 2023
Cited alongside, same era.
Multi-target multiplicity: Flexibility and fairness in target specification under resource constraints
Jamelle Watson-Daniels, Solon Barocas, Jake M Hofman, and Alexandra Chouldechova · 2023
Cited alongside, same era.
Closest in time.
A.I. has a measurement problem
Kevin Roose · 2024
Closest in time.
Position: Measure dataset diversity, don’t just claim it
Dora Zhao, Jerone TA Andrews, Orestis Papakyriakopoulos, and Alice Xiang · 2024
Closest in time.