Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have been observed to encode and perpetuate harmful associations present in the training data.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Crows-pairs: A challenge dataset for measuring social biases in masked language models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel Bowman. 2020 · 1967
Earlier work this paper cites.
A model of (often mixed) stereotype content: competence and warmth respectively follow from perceived status and competition
Susan T Fiske, Amy JC Cuddy, Peter Glick, and Jun Xu. 2002 · 2002
Earlier work this paper cites.
The bias map: behaviors from intergroup affect and stereotypes
Amy JC Cuddy, Susan T Fiske, and Peter Glick. 2007 · 2007
Earlier work this paper cites.
Facets of the fundamental content dimensions: Agency with competence and assertiveness—communion with warmth and morality
Andrea E Abele, Nicole Hauke, Kim Peters, Eva Louvet, Aleksandra Szymkow, and Yanping Duan. 2016 · 2016
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. 2016 · 2016
Earlier work this paper cites.
The abc of stereotypes about groups: Agency/socioeconomic success, conservative–progressive beliefs, and communion
Alex Koch, Roland Imhoff, Ron Dotsch, Christian Unkelbach, and Hans Alves. 2016 · 2016
Earlier work this paper cites.
Stereotype content: Warmth and competence endure
Susan T Fiske. 2018 · 2018
Earlier work this paper cites.
Generating informative and diverse conversational responses via adversarial information maximization
Yizhe Zhang, Michel Galley, Jianfeng Gao, Zhe Gan, Xiujun Li, Chris Brockett, and Bill Dolan. 2018 · 2018
Earlier work this paper cites.
Texygen: A benchmarking platform for text generation models
Yaoming Zhu, Sidi Lu, Lei Zheng, Jiaxian Guo, Weinan Zhang, Jun Wang, and Yong Yu. 2018 · 2018
Earlier work this paper cites.
Text summarization with pretrained encoders
Yang Liu and Mirella Lapata. 2019 · 2019
Cited alongside, same era.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Cited alongside, same era.
The risk of racial bias in hate speech detection
Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A Smith. 2019 · 2019
Cited alongside, same era.
Keybert: Minimal keyword extraction with bert
Maarten Grootendorst. 2020 · 2020
Cited alongside, same era.
Dialogpt: Large-scale generative pre-training for conversational response generation
Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and William B Dolan. 2020 · 2020
Cited alongside, same era.
Large language models associate muslims with violence
Abubakar Abid, Maheen Farooqi, and James Zou. 2021 · 2021
Stereoset: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy. 2021 · 2021
Later among the works it cites.
Comprehensive stereotype content dictionaries using a semi-automated method
Gandalf Nicolas, Xuechunzi Bai, and Susan T Fiske. 2021 · 2021
Later among the works it cites.
Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp
Timo Schick, Sahana Udupa, and Hinrich Schütze. 2021 · 2021
Later among the works it cites.
Theory-grounded measurement of us social stereotypes in english language models
Yang Cao, Anna Sotnikova, Hal Daumé III, Rachel Rudinger, and Linda Zou. 2022 · 2022
Later among the works it cites.
Generated knowledge prompting for commonsense reasoning
Jiacheng Liu, Alisa Liu, Ximing Lu, Sean Welleck, Peter West, Ronan Le Bras, Yejin Choi, and Hannaneh Hajishirzi. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Stereotyping norwegian salmon: An inventory of pitfalls in fairness benchmark datasets
Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach. 2021 · 2021
Cited alongside, same era.
Bold: Dataset and metrics for measuring biases in open-ended language generation
Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruksachatkun, Kai-Wei Chang, and Rahul Gupta. 2021 · 2021
Cited alongside, same era.
Understanding and countering stereotypes: A computational approach to the stereotype content model
Kathleen C Fraser, Isar Nejadgholi, and Svetlana Kiritchenko. 2021 · 2021
Cited alongside, same era.
Bias out-of-the-box: An empirical analysis of intersectional occupational biases in popular generative language models
Hannah Rose Kirk, Yennie Jun, Filippo Volpin, Haider Iqbal, Elias Benussi, Frederic Dreyer, Aleksandar Shtedritski, and Yuki Asano. 2021 · 2021
Cited alongside, same era.
Gandalf Nicolas, Xuechunzi Bai, and Susan T Fiske. 2022 · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. 2022 · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022 · 2022
Later among the works it cites.
Measuring normative and descriptive biases in language models using census data
Samia Touileb, Lilja Øvrelid, and Erik Velldal. 2023 · 2023
Closest in time.