Stereotyping norwegian salmon: an inventory of pitfalls in fairness benchmark datasets
Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach. 2021 · 2021
Later among the works it cites.
What do bias measures measure?
Original
Sunipa Dev, Emily Sheng, Jieyu Zhao, Jiao Sun, Yu Hou, Mattie Sanseverino, Jiin Kim, Nanyun Peng, and Kai-Wei Chang. 2021 · 2021
Later among the works it cites.
Bold: Dataset and metrics for measuring biases in open-ended language generation
Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruksachatkun, Kai-Wei Chang, and Rahul Gupta. 2021 · 2021
Later among the works it cites.
Bidimensional leaderboards: Generate and evaluate language hand in hand
Original
Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Lavinia Dunagan, Jacob Morrison, Alexander R Fabbri, Yejin Choi, and Noah A Smith. 2021 · 2021
Later among the works it cites.
Alexa tells 10-year-old girl to touch live plug with penny
BBC News. 2021 · 2021
Later among the works it cites.
Honest: Measuring hurtful sentence completion in language models
Debora Nozza, Federico Bianchi, and Dirk Hovy. 2021 · 2021
Later among the works it cites.
Multitask prompted training enables zero-shot task generalization
Original
Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al. 2021 · 2021
Later among the works it cites.
They, them, theirs: Rewriting with gender-neutral english
Original
Tony Sun, Kellie Webster, Apu Shah, William Yang Wang, and Melvin Johnson. 2021 · 2021
Later among the works it cites.
Challenges in detoxifying language models
Original
Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang. 2021 · 2021
Later among the works it cites.