Fetching the paper…
Reading the bibliography…
Social bias is shaped by the accumulation of social perceptions towards targets across various demographic identities.
Transfertransfo: A transfer learning approach for neural network based conversational agents
Thomas Wolf, Victor Sanh, Julien Chaumond, and Clement Delangue. 2019 · 1901
Earlier work this paper cites.
Release strategies and the social impacts of language models
Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, et al. 2019 · 1908
Earlier work this paper cites.
The nature of prejudice
Gordon Willard Allport, Kenneth Clark, and Thomas Pettigrew. 1954 · 1954
Earlier work this paper cites.
CrowS-pairs: A challenge dataset for measuring social biases in masked language models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020 · 1967
Earlier work this paper cites.
The hostile media phenomenon: biased perception and perceptions of media bias in coverage of the beirut massacre
Robert P Vallone, Lee Ross, and Mark R Lepper. 1985 · 1985
Earlier work this paper cites.
Why study stereotype accuracy and inaccuracy?
Lee J Jussim, Clark R McCauley, and Yueh-Ting Lee. 1995 · 1995
Earlier work this paper cites.
Measuring individual differences in implicit cognition: the implicit association test
Anthony G. Greenwald, Debbie E. McGhee, and Jordan L. K. Schwartz. 1998 · 1998
Earlier work this paper cites.
Crowdsourcing and language studies: the new generation of linguistic data
Robert Munro, Steven Bethard, Victor Kuperman, Vicky Tzuyin Lai, Robin Melnick, Christopher Potts, Tyler Schnoebelen, and Harry Tily. 2010 · 2010
Earlier work this paper cites.
Social perception as induction and inference: An integrative model of intergroup differentiation, ingroup favoritism, and differential accuracy
Theresa E DiDonato, Johannes Ullrich, and Joachim I Krueger. 2011 · 2011
Earlier work this paper cites.
Intergroup consensus/disagreement in support of group-based hierarchy: an examination of socio-structural and psycho-cultural factors
IC Lee, Felicia Pratto, and Blair T Johnson. 2011 · 2011
Earlier work this paper cites.
Social Psychology: 11th Edition
David Myers. 2012 · 2012
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. 2016 · 2016
Earlier work this paper cites.
A persona-based neural conversation model
Jiwei Li, Michel Galley, Chris Brockett, Georgios Spithourakis, Jianfeng Gao, and Bill Dolan. 2016 · 2016
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Aylin Caliskan, Joanna J. Bryson, and Arvind Narayanan. 2017 · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim. 2017 · 2017
Earlier work this paper cites.
Prejudices in cultural contexts: Shared stereotypes (gender, age) versus variable stereotypes (race, ethnicity, religion)
Susan T Fiske. 2017 · 2017
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M. Bender and Batya Friedman. 2018 · 2018
Earlier work this paper cites.
Fairness in machine learning: Lessons from political philosophy
Reuben Binns. 2018 · 2018
Earlier work this paper cites.
Gender shades: Intersectional accuracy disparities in commercial gender classification
Joy Buolamwini and Timnit Gebru. 2018 · 2018
Earlier work this paper cites.
Gender bias in coreference resolution
Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. 2018 · 2018
Earlier work this paper cites.
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018 · 2018
Earlier work this paper cites.
On measuring social biases in sentence encoders
Chandler May, Alex Wang, Shikha Bordia, Samuel R. Bowman, and Rachel Rudinger. 2019 · 2019
Earlier work this paper cites.
The woman worked as a babysitter: On biases in language generation
Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. 2019 · 2019
Cited alongside, same era.
Language (technology) is power: A critical survey of “bias” in NLP
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020 · 2020
Cited alongside, same era.
On measuring and mitigating biased inferences of word embeddings
Sunipa Dev, Tao Li, Jeff M Phillips, and Vivek Srikumar. 2020 · 2020
Cited alongside, same era.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020 · 2020
Cited alongside, same era.
Filtering before iteratively referring for knowledge-grounded response selection in retrieval-based chatbots
Jia-Chen Gu, Zhenhua Ling, Quan Liu, Zhigang Chen, and Xiaodan Zhu. 2020 · 2020
Cited alongside, same era.
Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. 2023 · 2023
Later among the works it cites.
Equi-tuning: Group equivariant fine-tuning of pretrained models
Sourya Basu, Prasanna Sattigeri, Karthikeyan Natesan Ramamurthy, Vijil Chenthamarakshan, Kush R Varshney, Lav R Varshney, and Payel Das. 2023 · 2023
Later among the works it cites.
Stereotypes in chatgpt: An empirical study
Tony Busker, Sunil Choenni, and Mortaza Shoae Bargh. 2023 · 2023
Later among the works it cites.
Marked personas: Using natural language prompts to measure stereotypes in language models
Myra Cheng, Esin Durmus, and Dan Jurafsky. 2023 · 2023
Later among the works it cites.
Toxicity in chatgpt: Analyzing persona-assigned language models
Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning to detect relevant contexts and knowledge for response selection in retrieval-based dialogue systems
Kai Hua, Zhiyuan Feng, Chongyang Tao, Rui Yan, and Lu Zhang. 2020 · 2020
Cited alongside, same era.
UNQOVERing stereotyping biases via underspecified questions
Tao Li, Daniel Khashabi, Tushar Khot, Ashish Sabharwal, and Vivek Srikumar. 2020 · 2020
Cited alongside, same era.
Assessing gender bias in machine translation: a case study with google translate
Marcelo OR Prates, Pedro H Avelar, and Luís C Lamb. 2020 · 2020
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Cited alongside, same era.
Bold: Dataset and metrics for measuring biases in open-ended language generation
Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruksachatkun, Kai-Wei Chang, and Rahul Gupta. 2021 · 2021
Cited alongside, same era.
Moral stories: Situated reasoning about norms, intents, actions, and their consequences
Denis Emelin, Ronan Le Bras, Jena D. Hwang, Maxwell Forbes, and Yejin Choi. 2021 · 2021
Cited alongside, same era.
Detecting emergent intersectional biases: Contextualized word embeddings contain a distribution of human-like biases
Wei Guo and Aylin Caliskan. 2021 · 2021
Cited alongside, same era.
ROBBIE: Robust bias evaluation of large generative language models
David Esiobu, Xiaoqing Tan, Saghar Hosseini, Megan Ung, Yuchen Zhang, Jude Fernandes, Jane Dwivedi-Yu, Eleonora Presani, Adina Williams, and Eric Smith. 2023 · 2023
Later among the works it cites.
Should chatgpt be biased? challenges and risks of bias in large language models
Emilio Ferrara. 2023 · 2023
Later among the works it cites.
Bias runs deep: Implicit reasoning biases in persona-assigned llms
Shashank Gupta, Vaishnavi Shrivastava, Ameet Deshpande, Ashwin Kalyan, Peter Clark, Ashish Sabharwal, and Tushar Khot. 2023 · 2023
Later among the works it cites.
Cbbq: A chinese bias benchmark dataset curated with human-ai collaboration for large language models
Yufei Huang and Deyi Xiong. 2023 · 2023
Later among the works it cites.
Better zero-shot reasoning with role-play prompting
Aobo Kong, Shiwan Zhao, Hao Chen, Qicheng Li, Yong Qin, Ruiqi Sun, and Xin Zhou. 2023 · 2023
Later among the works it cites.
Chatharuhi: Reviving anime character in reality via large language model
Cheng Li, Ziang Leng, Chenxi Yan, Junyi Shen, Hao Wang, Weishi MI, Yaying Fei, Xiaoyang Feng, Song Yan, HaoSheng Wang, et al. 2023 · 2023
Later among the works it cites.
OpenAI. 2023 · 2023
Later among the works it cites.
In-context impersonation reveals large language models’ strengths and biases
Leonard Salewski, Stephan Alaniz, Isabel Rio-Torto, Eric Schulz, and Zeynep Akata. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Later among the works it cites.
“kelly is a warm person, joseph is a role model”: Gender biases in LLM-generated reference letters
Yixin Wan, George Pu, Jiao Sun, Aparna Garimella, Kai-Wei Chang, and Nanyun Peng. 2023a · 2023
Later among the works it cites.
Are personalized stochastic parrots more dangerous? evaluating persona biases in dialogue systems
Yixin Wan, Jieyu Zhao, Aman Chadha, Nanyun Peng, and Kai-Wei Chang. 2023b · 2023
Later among the works it cites.
Expertprompting: Instructing large language models to be distinguished experts
Benfeng Xu, An Yang, Junyang Lin, Quan Wang, Chang Zhou, Yongdong Zhang, and Zhendong Mao. 2023 · 2023
Later among the works it cites.
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024 · 2024
Closest in time.
KoBBQ: Korean Bias Benchmark for Question Answering
Jiho Jin, Jiseon Kim, Nayeon Lee, Haneul Yoo, Alice Oh, and Hwaran Lee. 2024 · 2024
Closest in time.
A roadmap to pluralistic alignment
Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, et al. 2024 · 2024
Closest in time.
Prohibited employment policies/practices
U.S. Equal Employment Opportunity Commission. 2024 · 2024
Closest in time.
BBQ: A hand-built bias benchmark for question answering
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel Bowman. 2022 · 2086
Closest in time.