Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have shown remarkable advances in language generation and understanding but are also prone to exhibiting harmful social biases.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Plug and play language models: A simple approach to controlled text generation
Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. 2019 · 1912
Earlier work this paper cites.
Linguistic intergroup bias: Stereotype perpetuation through language
Anne Maass. 1999 · 1999
Earlier work this paper cites.
Academic race stereotypes, academic self-concept, and racial centrality in african american youth
Ndidi A Okeke, Lionel C Howard, Beth Kurtz-Costes, and Stephanie J Rowley. 2009 · 2009
Earlier work this paper cites.
Measuring and reducing gendered correlations in pre-trained models
Kellie Webster, Xuezhi Wang, Ian Tenney, Alex Beutel, Emily Pitler, Ellie Pavlick, Jilin Chen, Ed Chi, and Slav Petrov. 2020 · 2010
Earlier work this paper cites.
How stereotypes are shared through language: a review and introduction of the social categories and stereotypes communication (scsc) framework
Camiel J Beukeboom and Christian Burgers. 2019 · 2019
Earlier work this paper cites.
Reducing gender bias in word-level language models with a gender-equalizing loss function
Yusu Qian, Urwa Muaz, Ben Zhang, and Jae Won Hyun. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Earlier work this paper cites.
Counterfactual data augmentation for mitigating gender stereotypes in languages with rich morphology
Ran Zmigrod, Sabrina J. Mielke, Hanna Wallach, and Ryan Cotterell. 2019 · 2019
Earlier work this paper cites.
Language (technology) is power: A critical survey of “bias” in NLP
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020 · 2020
Earlier work this paper cites.
Queens are powerful too: Mitigating gender bias in dialogue generation
Emily Dinan, Angela Fan, Adina Williams, Jack Urbanek, Douwe Kiela, and Jason Weston. 2020 · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020 · 2020
Earlier work this paper cites.
Social biases in NLP models as barriers for persons with disabilities
Ben Hutchinson, Vinodkumar Prabhakaran, Emily Denton, Kellie Webster, Yu Zhong, and Stephen Denuyl. 2020 · 2020
Earlier work this paper cites.
Mitigating gender bias amplification in distribution by posterior regularization
Shengyu Jia, Tao Meng, Jieyu Zhao, and Kai-Wei Chang. 2020 · 2020
Earlier work this paper cites.
Does gender matter? towards fairness in dialogue systems
Haochen Liu, Jamell Dacon, Wenqi Fan, Hui Liu, Zitao Liu, and Jiliang Tang. 2020 · 2020
Earlier work this paper cites.
Gender bias in neural natural language processing
Kaiji Lu, Piotr Mardziel, Fangjing Wu, Preetam Amancharla, and Anupam Datta. 2020 · 2020
Earlier work this paper cites.
Towards Controllable Biases in Language Generation
Emily Sheng, Kai-Wei Chang, Prem Natarajan, and Nanyun Peng. 2020 · 2020
Earlier work this paper cites.
Persistent anti-muslim bias in large language models
Abubakar Abid, Maheen Farooqi, and James Zou. 2021 · 2021
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Earlier work this paper cites.
FairFil: Contrastive neural debiasing method for pretrained text encoders
Pengyu Cheng, Weituo Hao, Siyang Yuan, Shijing Si, and Lawrence Carin. 2021 · 2021
Earlier work this paper cites.
He is very intelligent, she is very beautiful? on mitigating social biases in language modelling and generation
Aparna Garimella, Akhash Amarnath, Kiran Kumar, Akash Pramod Yalla, N Anandhavelu, Niyati Chhaya, and Balaji Vasan Srinivasan. 2021 · 2021
Earlier work this paper cites.
Generating gender augmented data for NLP
Nishtha Jain, Maja Popović, Declan Groves, and Eva Vanmassenhove. 2021 · 2021
Earlier work this paper cites.
Debiasing pre-trained contextualised embeddings
Masahiro Kaneko and Danushka Bollegala. 2021 · 2021
Earlier work this paper cites.
GeDi: Generative discriminator guided sequence generation
Ben Krause, Akhilesh Deepak Gotmare, Bryan McCann, Nitish Shirish Keskar, Shafiq Joty, Richard Socher, and Nazneen Fatema Rajani. 2021 · 2021
Cited alongside, same era.
DExperts: Decoding-time controlled text generation with experts and anti-experts
Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A. Smith, and Yejin Choi. 2021 · 2021
Cited alongside, same era.
Prompt programming for large language models: Beyond the few-shot paradigm
Laria Reynolds and Kyle McDonell. 2021 · 2021
Cited alongside, same era.
Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp
Timo Schick, Sahana Udupa, and Hinrich Schütze. 2021 · 2021
Cited alongside, same era.
“Nice try, kiddo”: Investigating ad hominems in dialogue responses
Emily Sheng, Kai-Wei Chang, Prem Natarajan, and Nanyun Peng. 2021a · 2021
Cited alongside, same era.
Perturbation augmentation for fairer NLP
Rebecca Qian, Candace Ross, Jude Fernandes, Eric Michael Smith, Douwe Kiela, and Adina Williams. 2022 · 2022
Later among the works it cites.
First the worst: Finding better gender translations during beam search
Danielle Saunders, Rosie Sallis, and Bill Byrne. 2022 · 2022
Later among the works it cites.
Text style transfer for bias mitigation using masked language modeling
Ewoenam Kwaku Tokpo and Toon Calders. 2022 · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022 · 2022
Later among the works it cites.
Taxonomy of risks posed by language models
Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, Courtney Biles, Sasha Brown, Zac Kenton, Will Hawkins, Tom Stepleton, Abeba Birhane, Lisa Anne Hendricks, Laura Rimell, William Isaac, Julia Haas, Sean Legassick, Geoffrey Irving, and Iason Gabriel. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tony Sun, Kellie Webster, Apu Shah, William Yang Wang, and Melvin Johnson. 2021 · 2021
Cited alongside, same era.
NeuTral Rewriter: A rule-based and neural approach to automatic rewriting into gender neutral alternatives
Eva Vanmassenhove, Chris Emmery, and Dimitar Shterionov. 2021 · 2021
Cited alongside, same era.
Calibrate before use: Improving few-shot performance of language models
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021 · 2021
Cited alongside, same era.
Entropy-based attention regularization frees unintended bias mitigation from lists
Giuseppe Attanasio, Debora Nozza, Dirk Hovy, and Elena Baralis. 2022 · 2022
Cited alongside, same era.
Looking for a handsome carpenter! debiasing GPT-3 job advertisements
Conrad Borchers, Dalia Gala, Benjamin Gilburt, Eduard Oravkin, Wilfried Bounsi, Yuki M Asano, and Hannah Kirk. 2022 · 2022
Cited alongside, same era.
Debiasing Pretrained Text Encoders by Paying Attention to Paying Attention
Yacine Gaci, Boualem Benattallah, Fabio Casati, and Khalid Benabdeslem. 2022 · 2022
Cited alongside, same era.
Demographic-aware language model fine-tuning as a bias mitigation technique
Aparna Garimella, Rada Mihalcea, and Akhash Amarnath. 2022 · 2022
Cited alongside, same era.
Lichang Chen, Jiuhai Chen, Tom Goldstein, Heng Huang, and Tianyi Zhou. 2023 · 2023
Later among the works it cites.
Queer people are people first: Deconstructing sexual identity stereotypes in large language models
Harnoor Dhingra, Preetiha Jayashanker, Sayali Moghe, and Emma Strubell. 2023 · 2023
Later among the works it cites.
Improving gender fairness of pre-trained language models without catastrophic forgetting
Zahra Fatemi, Chen Xing, Wenhao Liu, and Caimming Xiong. 2023 · 2023
Later among the works it cites.
Gender-tuning: Empowering fine-tuning for debiasing pre-trained language models
Somayeh Ghanbarzadeh, Yan Huang, Hamid Palangi, Radames Cruz Moreno, and Hamed Khanpour. 2023 · 2023
Later among the works it cites.
Llm self defense: By self examination, llms know they are being tricked
Alec Helbling, Mansi Phute, Matthew Hull, and Duen Horng Chau. 2023 · 2023
Later among the works it cites.
Yunqi Li and Yongfeng Zhang. 2023 · 2023
Later among the works it cites.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023 · 2023
Later among the works it cites.
Using in-context learning to improve dialogue safety
Nicholas Meade, Spandana Gella, Devamanyu Hazarika, Prakhar Gupta, Di Jin, Siva Reddy, Yang Liu, and Dilek Hakkani-Tür. 2023 · 2023
Later among the works it cites.
Bias against 93 stigmatized groups in masked language models and downstream sentiment classification tasks
Katelyn Mei, Sonia Fereidooni, and Aylin Caliskan. 2023 · 2023
Later among the works it cites.
Nationality bias in text generation
Pranav Narayanan Venkit, Sanjana Gautam, Ruchi Panchanadikar, Ting-Hao Huang, and Shomir Wilson. 2023 · 2023
Later among the works it cites.
Never too late to learn: Regularizing gender bias in coreference resolution
SunYoung Park, Kyuri Choi, Haeun Yu, and Youngjoong Ko. 2023 · 2023
Later among the works it cites.
Compensatory debiasing for gender imbalances in language models
Tae-Jin Woo, Woo-Jeoung Nam, Yeong-Joon Ju, and Seong-Whan Lee. 2023 · 2023
Later among the works it cites.
Adept: A debiasing prompt framework
Ke Yang, Charles Yu, Yi R Fung, Manling Li, and Heng Ji. 2023 · 2023
Later among the works it cites.
Unlearning bias in language models by partitioning gradients
Charles Yu, Sullam Jeoung, Anish Kasi, Pengfei Yu, and Heng Ji. 2023 · 2023
Later among the works it cites.
Deep learning on a healthy data diet: Finding important examples for fairness
Abdelrahman Zayed, Prasanna Parthasarathi, Gonçalo Mordido, Hamid Palangi, Samira Shabanian, and Sarath Chandar. 2023 · 2023
Later among the works it cites.
Click: Controllable text generation with sequence likelihood contrastive learning
Chujie Zheng, Pei Ke, Zheng Zhang, and Minlie Huang. 2023 · 2023
Later among the works it cites.
BBQ: A hand-built bias benchmark for question answering
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel Bowman. 2022 · 2086
Closest in time.