Bbq: A hand-built bias benchmark for question answering
Original
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel R Bowman. 2021 · 2021
Later among the works it cites.
Few-shot instruction prompts for pretrained language models to detect social biases
Original
Shrimai Prabhumoye, Rafal Kocielnik, Mohammad Shoeybi, Anima Anandkumar, and Bryan Catanzaro. 2021 · 2021
Later among the works it cites.
Societal biases in language generation: Progress and challenges
Original
Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. 2021 · 2021
Later among the works it cites.
Worst of both worlds: Biases compound in pre-trained vision-and-language models
Original
Tejas Srinivasan and Yonatan Bisk. 2021 · 2021
Later among the works it cites.
On generalization in coreference resolution
Shubham Toshniwal, Patrick Xia, Sam Wiseman, Karen Livescu, and Kevin Gimpel. 2021 · 2021
Later among the works it cites.
Double perturbation: On the robustness of robustness and counterfactual bias evaluation
Chong Zhang, Jieyu Zhao, Huan Zhang, Kai-Wei Chang, and Cho-Jui Hsieh. 2021 · 2021
Later among the works it cites.
Ethical-advice taker: Do language models understand natural language interventions?
Original
Jieyu Zhao, Daniel Khashabi, Tushar Khot, Ashish Sabharwal, and Kai-Wei Chang. 2021 · 2021
Later among the works it cites.
Your fairness may vary: Pretrained language model fairness in toxic text classification
Ioana Baldini, Dennis Wei, Karthikeyan Natesan Ramamurthy, Moninder Singh, and Mikhail Yurochkin. 2022 · 2022
Closest in time.
Re-contextualizing fairness in NLP: The case of India
Shaily Bhatt, Sunipa Dev, Partha Talukdar, Shachi Dave, and Vinodkumar Prabhakaran. 2022 · 2022
Closest in time.
On the intrinsic and extrinsic fairness evaluation metrics for contextualized language representations
Yang Trista Cao, Yada Pruksachatkun, Kai-Wei Chang, Rahul Gupta, Varun Kumar, Jwala Dhamala, and Aram Galstyan. 2022 · 2022
Closest in time.
Palm: Scaling language modeling with pathways
Original
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam M. Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Benton C. Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier García, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Díaz, Orhan Firat, Michele Catasta, Jason Wei, Kathleen S. Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022 · 2022
Closest in time.
WANLI: Worker and AI Collaboration for Natural Language Inference Dataset Creation
Original
Alisa Liu, Swabha Swayamdipta, Noah A Smith, and Yejin Choi. 2022 · 2022
Closest in time.
Large pre-trained language models contain human-like biases of what is right and wrong to do
Original
Patrick Schramowski, Cigdem Turan, Nico Andersen, Constantin A Rothkopf, and Kristian Kersting. 2022 · 2022
Closest in time.
Quantifying social biases using templates is unreliable
Original
Preethi Seshadri, Pouya Pezeshkpour, and Sameer Singh. 2022 · 2022
Closest in time.
LaMDA: Language Models for Dialog Applications
Original
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. 2022 · 2022
Closest in time.
This prompt is measuring< mask>: Evaluating bias evaluation in language models
Original
Seraphina Goldfarb-Tarrant, Eddie Ungless, Esma Balkir, and Su Lin Blodgett. 2023 · 2023
Closest in time.