Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs), now used daily by millions, can encode societal biases, exposing their users to representational harms.
Mitigating Gender Bias in Natural Language Processing: Literature Review
Tony Sun, Andrew Gaut, Shirlyn Tang, Yuxin Huang, Mai ElSherief, Jieyu Zhao, Diba Mirza, Elizabeth Belding, Kai-Wei Chang, and William Yang Wang. 2019 · 1906
Earlier work this paper cites.
Unsupervised Cross-lingual Representation Learning at Scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 1911
Earlier work this paper cites.
Divide and rule: British policy in indian history
Neil Stewart. 1951 · 1951
Earlier work this paper cites.
Status evaluation in the Hindu caste system
Henry Noel Cochran Stevenson. 1954 · 1954
Earlier work this paper cites.
The mental pictures of six Hindu caste groups about each other as reflected in verbal stereotypes
R Rath and NC Sircar. 1960 · 1960
Earlier work this paper cites.
Exploration in caste stereotypes
Gopal Sharan Sinha and Ramesh Chandra Sinha. 1967 · 1967
Earlier work this paper cites.
The Pre-history of ‘; Communalism’? Religious Conflict in India, 1700–1860
Christopher A Bayly. 1985 · 1985
Earlier work this paper cites.
The crime of punishment: Racial and gender disparities in the use of corporal punishment in US public schools
James F Gregory. 1995 · 1995
Earlier work this paper cites.
Inscribing the other, inscribing the self: Hindu-Muslim identities in pre-colonial India
Cynthia Talbot. 1995 · 1995
Earlier work this paper cites.
Racial and gender biases in magazine advertising: A content-analytic study
Scott Plous and Dominique Neptune. 1997 · 1997
Earlier work this paper cites.
‘Race’, religion and riots: The ‘racialization’of communal identity and conflict in India
Zaheer Baber. 2004 · 2004
Earlier work this paper cites.
Language Models are Few-Shot Learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2005
Earlier work this paper cites.
History, emotions and hetero-referential representations in inter-group conflict: The example of Hindu-Muslim relations in India
Ragini Sen and Wolfgang Wagner. 2005 · 2005
Earlier work this paper cites.
(MIS) Representing the Dalit Woman: Reification of Caste and Gender Stereotypes in the Hindi Didactic Literature of Colonial India
Charu Gupta. 2008a · 2008
Earlier work this paper cites.
(MIS) Representing the Dalit Woman: Reification of Caste and Gender Stereotypes in the Hindi Didactic Literature of Colonial India
Charu Gupta. 2008b · 2008
Earlier work this paper cites.
Is Caste Intrinsic to Hinduism?
Anantanand Rambachan. 2008 · 2008
Earlier work this paper cites.
Religion, socio-economic backwardness & discrimination: The case of Indian Muslims
Rowena Robinson. 2008 · 2008
Earlier work this paper cites.
The political logic of ethnic violence: The anti-Muslim pogrom in Gujarat, 2002
Raheel Dhattiwala and Michael Biggs. 2012 · 2012
Earlier work this paper cites.
Religion insulates ingroup evaluations: The development of intergroup attitudes in India
Yarrow Dunham, Mahesh Srinivasan, Ron Dotsch, and David Barner. 2014 · 2014
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. 2016 · 2016
Earlier work this paper cites.
Changes in racial and gender inequality since 1970
C Matthew Snipp and Sin Yi Cheung. 2016 · 2016
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Caste system, Dalitization and its implications in contemporary India
Selvin Raj Gnana. 2018 · 2018
Earlier work this paper cites.
Hate speech detection from code-mixed hindi-english tweets using deep learning models
Satyajit Kamble and Aditya Joshi. 2018 · 2018
Earlier work this paper cites.
Racial bias in hate speech and abusive language detection datasets
Thomas Davidson, Debasmita Bhattacharya, and Ingmar Weber. 2019 · 2019
Cited alongside, same era.
Counterfactual fairness in text classification through robustness. In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society . 219–226
Sahaj Garg, Vincent Perot, Nicole Limtiaco, Ankur Taly, Ed H Chi, and Alex Beutel. 2019 · 2019
Cited alongside, same era.
End-to-end bias mitigation by modelling biases in corpora
Rabeeh Karimi Mahabadi, Yonatan Belinkov, and James Henderson. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Language (technology) is power: A critical survey of" bias" in nlp
You reap what you sow: On the challenges of bias evaluation under multilingual settings. In Proceedings of BigScience Episode# 5–Workshop on Challenges & Perspectives in Creating Large Language Models . 26–41
Zeerak Talat, Aurélie Névéol, Stella Biderman, Miruna Clinciu, Manan Dey, Shayne Longpre, Sasha Luccioni, Maraim Masoud, Margaret Mitchell, Dragomir Radev, et al · 2022
Later among the works it cites.
Attitudes about Caste
Pew Research Center. 2021 · 2023
Closest in time.
Building Socio-culturally Inclusive Stereotype Resources with Community Engagement
Sunipa Dev, Jaya Goyal, Dinesh Tewari, Shachi Dave, and Vinodkumar Prabhakaran. 2023 · 2023
Closest in time.
Queer People are People First: Deconstructing Sexual Identity Stereotypes in Large Language Models
Harnoor Dhingra, Preetiha Jayashanker, Sayali Moghe, and Emma Strubell. 2023 · 2023
Closest in time.
Raft: Reward ranked finetuning for generative foundation model alignment
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020 · 2020
Cited alongside, same era.
On transferability of bias mitigation effects in language model fine-tuning
Xisen Jin, Francesco Barbieri, Brendan Kennedy, Aida Mostafazadeh Davani, Leonardo Neves, and Xiang Ren. 2020 · 2020
Cited alongside, same era.
BR Ambedkar on caste and land relations in India
Awanish Kumar. 2020 · 2020
Cited alongside, same era.
StereoSet: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy. 2020 · 2020
Cited alongside, same era.
CrowS-pairs: A challenge dataset for measuring social biases in masked language models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R Bowman. 2020 · 2020
Cited alongside, same era.
Social inequalities, fundamental inequities, and recurring of the digital divide: Insights from India
Nidhi Tewathia, Anant Kamath, and P Vigneswara Ilavarasan. 2020 · 2020
Cited alongside, same era.
TIMUR’S INVASION OF INDIA
Ranjodh Jamwal. 2021 · 2021
Cited alongside, same era.
Bias out-of-the-box: An empirical analysis of intersectional occupational biases in popular generative language models
Hannah Rose Kirk, Yennie Jun, Filippo Volpin, Haider Iqbal, Elias Benussi, Frederic Dreyer, Aleksandar Shtedritski, and Yuki Asano. 2021 · 2021
Cited alongside, same era.
Hanze Dong, Wei Xiong, Deepanshu Goyal, Yihan Zhang, Winnie Chow, Rui Pan, Shizhe Diao, Jipeng Zhang, Kashun Shum, and Tong Zhang. 2023 · 2023
Closest in time.
Breaking the Bias: Gender Fairness in LLMs Using Prompt Engineering and In-Context Learning
Satyam Dwivedi, Sanjukta Ghosh, and Shivam Dwivedi. 2023 · 2023
Closest in time.
Winoqueer: A community-in-the-loop benchmark for anti-lgbtq+ bias in large language models
Virginia K Felkner, Ho-Chun Herbert Chang, Eugene Jang, and Jonathan May. 2023 · 2023
Closest in time.
Shangbin Feng, Chan Young Park, Yuhan Liu, and Yulia Tsvetkov. 2023 · 2023
Closest in time.
Decoder-only or encoder-decoder? interpreting language model as a regularized encoder-decoder
Zihao Fu, Wai Lam, Qian Yu, Anthony Man-Cho So, Shengding Hu, Zhiyuan Liu, and Nigel Collier. 2023 · 2023
Closest in time.
The capacity for moral self-correction in large language models
Deep Ganguli, Amanda Askell, Nicholas Schiefer, Thomas Liao, Kamilė Lukošiūtė, Anna Chen, Anna Goldie, Azalia Mirhoseini, Catherine Olsson, Danny Hernandez, et al · 2023
Closest in time.
ChatGPT sets record for fastest-growing user base - analyst note
Krystal Hu. 2023 · 2023
Closest in time.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Closest in time.
Hannah Rose Kirk, Andrew M Bean, Bertie Vidgen, Paul Röttger, and Scott A Hale. 2023 · 2023
Closest in time.
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023 · 2023
Closest in time.
A trip towards fairness: Bias and de-biasing in large language models
Leonardo Ranaldi, Elena Sofia Ruzzetti, Davide Venditti, Dario Onorati, and Fabio Massimo Zanzotto. 2023 · 2023
Closest in time.
Whose opinions do language models reflect?
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023 · 2023
Closest in time.
Vishesh Thakur. 2023 · 2023
Closest in time.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023 · 2023
Closest in time.
On evaluating and mitigating gender biases in multilingual settings
Aniket Vashishtha, Kabir Ahuja, and Sunayana Sitaram. 2023 · 2023
Closest in time.
Tianyu Yu, Yuan Yao, Haoye Zhang, Taiwen He, Yifeng Han, Ganqu Cui, Jinyi Hu, Zhiyuan Liu, Hai-Tao Zheng, Maosong Sun, et al · 2023
Closest in time.
ChatGPT exhibits gender and racial biases in acute coronary syndrome management
Angela Zhang, Mert Yuksekgonul, Joshua Guild, James Zou, and Joseph Wu. 2023 · 2023
Closest in time.
Steering LLMs Towards Unbiased Responses: A Causality-Guided Debiasing Framework
Jingling Li, Zeyu Tang, Xiaoyu Liu, Peter Spirtes, Kun Zhang, Liu Leqi, and Yang Liu. 2024 · 2024
Closest in time.
Chen Zheng, Ke Sun, Hang Wu, Chenguang Xi, and Xun Zhou. 2024 · 2024
Closest in time.