Fetching the paper…
Reading the bibliography…
Large language models (LLMs) trained on vast corpora suffer from inevitable stereotype biases.
CrowS-pairs: A challenge dataset for measuring social biases in masked language models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020 · 1967
Earlier work this paper cites.
Measuring and reducing gendered correlations in pre-trained models
Kellie Webster, Xuezhi Wang, Ian Tenney, Alex Beutel, Emily Pitler, Ellie Pavlick, Jilin Chen, Ed Chi, and Slav Petrov. 2020 · 2010
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018 · 2018
Earlier work this paper cites.
Counterfactual data augmentation for mitigating gender stereotypes in languages with rich morphology
Ran Zmigrod, Sabrina J. Mielke, Hanna Wallach, and Ryan Cotterell. 2019 · 2019
Earlier work this paper cites.
Unmasking contextual stereotypes: Measuring and mitigating BERT’s gender bias
Marion Bartl, Malvina Nissim, and Albert Gatt. 2020 · 2020
Earlier work this paper cites.
Queens are powerful too: Mitigating gender bias in dialogue generation
Emily Dinan, Angela Fan, Adina Williams, Jack Urbanek, Douwe Kiela, and Jason Weston. 2020 · 2020
Earlier work this paper cites.
Reducing sentiment bias in language models via counterfactual evaluation
Po-Sen Huang, Huan Zhang, Ray Jiang, Robert Stanforth, Johannes Welbl, Jack Rae, Vishal Maini, Dani Yogatama, and Pushmeet Kohli. 2020 · 2020
Earlier work this paper cites.
Does gender matter? towards fairness in dialogue systems
Haochen Liu, Jamell Dacon, Wenqi Fan, Hui Liu, Zitao Liu, and Jiliang Tang. 2020 · 2020
Earlier work this paper cites.
Stanza: A python natural language processing toolkit for many human languages
Peng Qi, Yuhao Zhang, Yuhui Zhang, Jason Bolton, and Christopher D Manning. 2020 · 2020
Earlier work this paper cites.
Modifying memories in transformer models
Chen Zhu, Ankit Singh Rawat, Manzil Zaheer, Srinadh Bhojanapalli, Daliang Li, Felix Yu, and Sanjiv Kumar. 2020 · 2020
Earlier work this paper cites.
Persistent anti-muslim bias in large language models
Abubakar Abid, Maheen Farooqi, and James Zou. 2021 · 2021
Earlier work this paper cites.
Editing factual knowledge in language models
Nicola De Cao, Wilker Aziz, and Ivan Titov. 2021 · 2021
Earlier work this paper cites.
Detect and perturb: Neutral rewriting of biased and sensitive text via gradient-based decoding
Zexue He, Bodhisattwa Prasad Majumder, and Julian McAuley. 2021 · 2021
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021 · 2021
Earlier work this paper cites.
Balancing out bias: Achieving fairness through balanced training
Xudong Han, Timothy Baldwin, and Trevor Cohn. 2022 · 2022
Earlier work this paper cites.
Controlling bias exposure for fair interpretable predictions
Zexue He, Yu Wang, Julian McAuley, and Bodhisattwa Prasad Majumder. 2022 · 2022
Earlier work this paper cites.
TruthfulQA: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022 · 2022
Earlier work this paper cites.
Interfair: Debiasing with natural language feedback for fair interpretable predictions
Bodhisattwa Prasad Majumder, Zexue He, and Julian McAuley. 2022 · 2022
Earlier work this paper cites.
Understanding stereotypes in language models: Towards robust measurement and zero-shot debiasing
Justus Mattern, Zhijing Jin, Mrinmaya Sachan, Rada Mihalcea, and Bernhard Schölkopf. 2022 · 2022
Cited alongside, same era.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Cited alongside, same era.
Perturbation augmentation for fairer NLP
Rebecca Qian, Candace Ross, Jude Fernandes, Eric Michael Smith, Douwe Kiela, and Adina Williams. 2022a · 2022
Cited alongside, same era.
Perturbation augmentation for fairer NLP
Rebecca Qian, Candace Ross, Jude Fernandes, Eric Michael Smith, Douwe Kiela, and Adina Williams. 2022b · 2022
Cited alongside, same era.
Critic-guided decoding for controlled text generation
Minbeom Kim, Hwanhee Lee, Kang Min Yoo, Joonsuk Park, Hwaran Lee, and Kyomin Jung. 2023 · 2023
Later among the works it cites.
Prompt tuning pushes farther, contrastive learning pulls closer: A two-stage approach to mitigate social biases
Yingji Li, Mengnan Du, Xin Wang, and Ying Wang. 2023 · 2023
Later among the works it cites.
Shengyu Mao, Ningyu Zhang, Xiaohan Wang, Mengru Wang, Yunzhi Yao, Yong Jiang, Pengjun Xie, Fei Huang, and Huajun Chen. 2023 · 2023
Later among the works it cites.
Using in-context learning to improve dialogue safety
Nicholas Meade, Spandana Gella, Devamanyu Hazarika, Prakhar Gupta, Di Jin, Siva Reddy, Yang Liu, and Dilek Hakkani-Tür. 2023 · 2023
Later among the works it cites.
Nationality bias in text generation
Pranav Narayanan Venkit, Sanjana Gautam, Ruchi Panchanadikar, Ting-Hao Huang, and Shomir Wilson. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
First the worst: Finding better gender translations during beam search
Danielle Saunders, Rosie Sallis, and Bill Byrne. 2022 · 2022
Cited alongside, same era.
Text style transfer for bias mitigation using masked language modeling
Ewoenam Kwaku Tokpo and Toon Calders. 2022 · 2022
Cited alongside, same era.
DUnE: Dataset for unified editing
Afra Akyürek, Eric Pan, Garry Kuwanto, and Derry Wijaya. 2023 · 2023
Cited alongside, same era.
Increasing diversity while maintaining accuracy: Text data generation with large language models and human interventions
John Chung, Ece Kamar, and Saleema Amershi. 2023 · 2023
Cited alongside, same era.
Queer people are people first: Deconstructing sexual identity stereotypes in large language models
Harnoor Dhingra, Preetiha Jayashanker, Sayali Moghe, and Emma Strubell. 2023 · 2023
Cited alongside, same era.
Palm-e: An embodied multimodal language model
Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint, Klaus Greff, Andy Zeng, Igor Mordatch, and Pete Florence. 2023 · 2023
Cited alongside, same era.
Improving gender fairness of pre-trained language models without catastrophic forgetting
Zahra Fatemi, Chen Xing, Wenhao Liu, and Caimming Xiong. 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
Never too late to learn: Regularizing gender bias in coreference resolution
SunYoung Park, Kyuri Choi, Haeun Yu, and Youngjoong Ko. 2023 · 2023
Later among the works it cites.
Layered bias: Interpreting bias in pretrained large language models
Nirmalendu Prakash and Roy Ka-Wei Lee. 2023 · 2023
Later among the works it cites.
A trip towards fairness: Bias and de-biasing in large language models
Leonardo Ranaldi, Elena Sofia Ruzzetti, Davide Venditti, Dario Onorati, and Fabio Massimo Zanzotto. 2023 · 2023
Later among the works it cites.
Compensatory debiasing for gender imbalances in language models
Tae-Jin Woo, Woo-Jeoung Nam, Yeong-Joon Ju, and Seong-Whan Lee. 2023 · 2023
Later among the works it cites.
Adept: A debiasing prompt framework
Ke Yang, Charles Yu, Yi R Fung, Manling Li, and Heng Ji. 2023 · 2023
Later among the works it cites.
Unlearning bias in language models by partitioning gradients
Charles Yu, Sullam Jeoung, Anish Kasi, Pengfei Yu, and Heng Ji. 2023 · 2023
Later among the works it cites.
Can we edit factual knowledge by in-context learning?
Ce Zheng, Lei Li, Qingxiu Dong, Yuxuan Fan, Zhiyong Wu, Jingjing Xu, and Baobao Chang. 2023 · 2023
Later among the works it cites.
Causal-debias: Unifying debiasing in pretrained language models and fine-tuning via causal invariant learning
Fan Zhou, Yuzhou Mao, Liu Yu, Yi Yang, and Ting Zhong. 2023 · 2023
Later among the works it cites.
Representation engineering: A top-down approach to ai transparency
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, and Dan Hendrycks. 2023 · 2023
Later among the works it cites.
Debiasing algorithm through model adaptation
Tomasz Limisiewicz, David Mareček, and Tomáš Musil. 2024 · 2024
Closest in time.
Massive editing for large language models via meta learning
Chenmien Tan, Ge Zhang, and Jie Fu. 2024 · 2024
Closest in time.
A comprehensive study of knowledge editing for large language models
Ningyu Zhang, Yunzhi Yao, Bozhong Tian, Peng Wang, Shumin Deng, Mengru Wang, Zekun Xi, Shengyu Mao, Jintian Zhang, Yuansheng Ni, Siyuan Cheng, Ziwen Xu, Xin Xu, Jia-Chen Gu, Yong Jiang, Pengjun Xie, Fei Huang, Lei Liang, Zhiqiang Zhang, Xiaowei Zhu, Jun Zhou, and Huajun Chen. 2024 · 2024
Closest in time.