Fetching the paper…
Reading the bibliography…
Sociodemographic bias in language models (LMs) has the potential for harm when deployed in real-world settings.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
Bad Seeds: Evaluating Lexical Methods for Bias Measurement
Maria Antoniak and David Mimno. 2021 · 1904
Earlier work this paper cites.
Probabilistic bias mitigation in word embeddings
Hailey James and David Alvarez-Melis. 2019 · 1910
Earlier work this paper cites.
Fairness without demographics in repeated loss minimization
Tatsunori Hashimoto, Megha Srivastava, Hongseok Namkoong, and Percy Liang. 2018 · 1938
Earlier work this paper cites.
Intrinsic Bias Metrics Do Not Correlate with Application Bias
Seraphina Goldfarb-Tarrant, Rebecca Marchant, Ricardo Muñoz Sánchez, Mugdha Pandya, and Adam Lopez. 2021 · 1940
Earlier work this paper cites.
CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020 · 1967
Earlier work this paper cites.
Logic and conversation
H. P. Grice. 1975 · 1975
Earlier work this paper cites.
How sensitive are translation systems to extra contexts? mitigating gender bias in neural machine translation models through relevant contexts
Shanya Sharma, Manan Dey, and Koustuv Sinha. 2022 · 1984
Earlier work this paper cites.
Automatically identifying gender issues in machine translation using perturbations
Hila Gonen and Kellie Webster. 2020 · 1995
Earlier work this paper cites.
Measuring individual differences in implicit cognition: the implicit association test
Anthony G Greenwald, Debbie E McGhee, and Jordan LK Schwartz. 1998 · 1998
Earlier work this paper cites.
PEFTDebias : Capturing debiasing information using PEFTs
Sumit Agarwal, Aditya Veerubhotla, and Srijan Bansal. 2023 · 2000
Earlier work this paper cites.
Grammatical gender associations outweigh topical gender bias in crosslinguistic word embeddings
Katherine McCurdy and Oguz Serbetci. 2020 · 2005
Earlier work this paper cites.
A meta-analytic test of intergroup contact theory
Thomas F Pettigrew and Linda R Tropp. 2006 · 2006
Earlier work this paper cites.
Warmth and competence as universal dimensions of social perception: The stereotype content model and the bias map
Amy JC Cuddy, Susan T Fiske, and Peter Glick. 2008 · 2008
Earlier work this paper cites.
Measuring and Reducing Gendered Correlations in Pre-trained Models
Kellie Webster, Xuezhi Wang, Ian Tenney, Alex Beutel, Emily Pitler, Ellie Pavlick, Jilin Chen, Ed Chi, and Slav Petrov. 2021 · 2010
Earlier work this paper cites.
Thinking, fast and slow
Daniel Kahneman. 2011 · 2011
Earlier work this paper cites.
Fairness through awareness
Cynthia Dwork, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Richard Zemel. 2012 · 2012
Earlier work this paper cites.
Tagging performance correlates with author age
Dirk Hovy and Anders Søgaard. 2015 · 2015
Earlier work this paper cites.
Demographic dialectal variation in social media: A case study of African-American English
Su Lin Blodgett, Lisa Green, and Brendan O’Connor. 2016 · 2016
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. 2016 · 2016
Earlier work this paper cites.
The problem with bias: Allocative versus representational harms in machine learning
Solon Barocas, Kate Crawford, Aaron Shapiro, and Hanna Wallach. 2017 · 2017
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017 · 2017
Earlier work this paper cites.
The trouble with bias
Kate Crawford. 2017 · 2017
Earlier work this paper cites.
Counterfactual fairness
Matt J Kusner, Joshua Loftus, Chris Russell, and Ricardo Silva. 2017 · 2017
Earlier work this paper cites.
Gender as a variable in natural-language processing: Ethical considerations
Brian Larson. 2017 · 2017
Earlier work this paper cites.
Social bias in elicited natural language inferences
Rachel Rudinger, Chandler May, and Benjamin Van Durme. 2017 · 2017
Earlier work this paper cites.
Men also like shopping: Reducing gender bias amplification using corpus-level constraints
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2017 · 2017
Earlier work this paper cites.
Addressing age-related bias in sentiment analysis
Mark Diaz, Isaac Johnson, Amanda Lazar, Anne Marie Piper, and Darren Gergle. 2018 · 2018
Earlier work this paper cites.
Measuring and Mitigating Unintended Bias in Text Classification
Lucas Dixon, John Li, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2018 · 2018
Earlier work this paper cites.
Word embeddings quantify 100 years of gender and ethnic stereotypes
Nikhil Garg, Londa Schiebinger, Dan Jurafsky, and James Zou. 2018 · 2018
Earlier work this paper cites.
Gender Recognition or Gender Reductionism?: The Social Implications of Embedded Gender Recognition Systems
Foad Hamidi, Morgan Klaus Scheuerman, and Stacy M. Branham. 2018 · 2018
Earlier work this paper cites.
Women also snowboard: Overcoming bias in captioning models
Lisa Anne Hendricks, Kaylee Burns, Kate Saenko, Trevor Darrell, and Anna Rohrbach. 2018 · 2018
Earlier work this paper cites.
Examining gender and race bias in two hundred sentiment analysis systems
Svetlana Kiritchenko and Saif Mohammad. 2018 · 2018
Earlier work this paper cites.
Reducing Gender Bias in Abusive Language Detection
Ji Ho Park, Jamin Shin, and Pascale Fung. 2018 · 2018
Earlier work this paper cites.
Gender Bias in Coreference Resolution
Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. 2018 · 2018
Earlier work this paper cites.
Darling or babygirl? investigating stylistic bias in sentiment analysis
Judy Hanwen Shen, Lauren Fratamico, Iyad Rahwan, and Alexander M Rush. 2018 · 2018
Earlier work this paper cites.
Biased embeddings from wild data: Measuring, understanding and removing
Adam Sutton, Thomas Lansdall-Welfare, and Nello Cristianini. 2018 · 2018
Earlier work this paper cites.
Mind the GAP: A balanced corpus of gendered ambiguous pronouns
Kellie Webster, Marta Recasens, Vera Axelrod, and Jason Baldridge. 2018 · 2018
Earlier work this paper cites.
Equality of opportunity in classification: A causal approach
Junzhe Zhang and Elias Bareinboim. 2018 · 2018
Earlier work this paper cites.
Gender Bias in Coreference Resolution: Evaluation and Debiasing Methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018a · 2018
Earlier work this paper cites.
Learning gender-neutral word embeddings
Jieyu Zhao, Yichao Zhou, Zeyu Li, Wei Wang, and Kai-Wei Chang. 2018b · 2018
Earlier work this paper cites.
Word embeddings (also) encode human personality stereotypes
Oshin Agarwal, Funda Durupınar, Norman I. Badler, and Ani Nenkova. 2019 · 2019
Earlier work this paper cites.
Differential privacy has disparate impact on model accuracy
Eugene Bagdasaryan, Omid Poursaeed, and Vitaly Shmatikov. 2019 · 2019
Earlier work this paper cites.
Evaluating the underlying gender bias in contextualized word embeddings
Christine Basta, Marta R. Costa-jussà, and Noe Casas. 2019 · 2019
Earlier work this paper cites.
A typology of ethical risks in language technology with an eye towards where transparent documentation can help. the future of artificial intelligence: Language
Emily M Bender. 2019 · 2019
Earlier work this paper cites.
Good secretaries, bad truck drivers? Occupational gender stereotypes in sentiment analysis
Jayadev Bhaskaran and Isha Bhallamudi. 2019 · 2019
Earlier work this paper cites.
Identifying and reducing gender bias in word-level language models
Shikha Bordia and Samuel R. Bowman. 2019 · 2019
Earlier work this paper cites.
Nuanced metrics for measuring unintended bias with real data for text classification
Daniel Borkan, Lucas Dixon, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2019 · 2019
Earlier work this paper cites.
Measuring gender bias in word embeddings across domains and discovering new gender bias word categories
Kaytlin Chaloner and Alfredo Maldonado. 2019 · 2019
Earlier work this paper cites.
On measuring gender bias in translation of gender-neutral pronouns
Won Ik Cho, Ji Won Kim, Seok Min Kim, and Nam Soo Kim. 2019 · 2019
Earlier work this paper cites.
Racial bias in hate speech and abusive language detection datasets
Thomas Davidson, Debasmita Bhattacharya, and Ingmar Weber. 2019 · 2019
Earlier work this paper cites.
Bias in Bios: A Case Study of Semantic Representation Bias in a High-Stakes Setting
Maria De-Arteaga, Alexey Romanov, Hanna Wallach, Jennifer Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Geyik, Krishnaram Kenthapadi, and Adam Tauman Kalai. 2019 · 2019
Earlier work this paper cites.
Attenuating bias in word vectors
Sunipa Dev and Jeff Phillips. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Equalizing gender bias in neural machine translation with word embeddings techniques
Joel Escudé Font and Marta R. Costa-jussà. 2019 · 2019
Earlier work this paper cites.
Understanding Undesirable Word Embedding Associations
Kawin Ethayarajh, David Duvenaud, and Graeme Hirst. 2019 · 2019
Earlier work this paper cites.
Relating word embedding gender biases to gender gaps: A cross-cultural analysis
Scott Friedman, Sonja Schmer-Galunder, Anthony Chen, and Jeffrey Rye. 2019 · 2019
Earlier work this paper cites.
Counterfactual fairness in text classification through robustness
Sahaj Garg, Vincent Perot, Nicole Limtiaco, Ankur Taly, Ed H. Chi, and Alex Beutel. 2019 · 2019
Earlier work this paper cites.
Women’s syntactic resilience and men’s grammatical luck: Gender-bias in part-of-speech tagging and dependency parsing
Aparna Garimella, Carmen Banea, Dirk Hovy, and Rada Mihalcea. 2019 · 2019
Earlier work this paper cites.
Lipstick on a pig: Debiasing methods cover up systematic gender biases in word embeddings but do not remove them
Hila Gonen and Yoav Goldberg. 2019 · 2019
Earlier work this paper cites.
Good” isn’t good enough
Ben Green. 2019 · 2019
Earlier work this paper cites.
deb2viz: Debiasing gender in word embedding data using subspace visualization
Enoch Opanin Gyamfi, Yunbo Rao, Miao Gou, and Yanhua Shao. 2020 · 2019
Earlier work this paper cites.
Unlearn dataset bias in natural language inference by fitting the residual
He He, Sheng Zha, and Haohan Wang. 2019 · 2019
Earlier work this paper cites.
50 years of test (un)fairness: Lessons for machine learning
Ben Hutchinson and Margaret Mitchell. 2019 · 2019
Earlier work this paper cites.
Gender-preserving debiasing for pre-trained word embeddings
Masahiro Kaneko and Danushka Bollegala. 2019 · 2019
Earlier work this paper cites.
Conceptor debiasing of word representations evaluated on WEAT
Saket Karve, Lyle Ungar, and João Sedoc. 2019 · 2019
Earlier work this paper cites.
Measuring bias in contextualized word representations
Keita Kurita, Nidhi Vyas, Ayush Pareek, Alan W Black, and Yulia Tsvetkov. 2019 · 2019
Earlier work this paper cites.
Are we consistently biased? multidimensional analysis of biases in distributional word vectors
Anne Lauscher and Goran Glavaš. 2019 · 2019
Earlier work this paper cites.
Black is to criminal as caucasian is to police: Detecting and removing multiclass bias in word embeddings
Thomas Manzini, Lim Yao Chong, Alan W Black, and Yulia Tsvetkov. 2019 · 2019
Earlier work this paper cites.
It’s all in the name: Mitigating gender bias with name-based counterfactual data substitution
Rowan Hall Maudslay, Hila Gonen, Ryan Cotterell, and Simone Teufel. 2019 · 2019
Earlier work this paper cites.
On measuring social biases in sentence encoders
Chandler May, Alex Wang, Shikha Bordia, Samuel R. Bowman, and Rachel Rudinger. 2019 · 2019
Earlier work this paper cites.
Equity Beyond Bias in Language Technologies for Education
Elijah Mayfield, Michael Madaio, Shrimai Prabhumoye, David Gerritsen, Brittany McLaughlin, Ezekiel Dixon-Román, and Alan W Black. 2019 · 2019
Earlier work this paper cites.
Disability, Bias, and AI
Mara Mills and Meredith Whittaker. 2019 · 2019
Earlier work this paper cites.
This thing called fairness: Disciplinary confusion realizing a value in technology
Deirdre K. Mulligan, Joshua A. Kroll, Nitin Kohli, and Richmond Y. Wong. 2019 · 2019
Earlier work this paper cites.
Unintended bias in misogyny detection
Debora Nozza, Claudia Volpetti, and Elisabetta Fersini. 2019 · 2019
Earlier work this paper cites.
Perturbation Sensitivity Analysis to Detect Unintended Model Biases
Vinodkumar Prabhakaran, Ben Hutchinson, and Margaret Mitchell. 2019 · 2019
Earlier work this paper cites.
Debiasing embeddings for reduced gender bias in text classification
Flavien Prost, Nithum Thain, and Tolga Bolukbasi. 2019 · 2019
Earlier work this paper cites.
Debiasing gender biased hindi words with word-embedding
Arun K. Pujari, Ansh Mittal, Anshuman Padhi, Anshul Jain, Mukesh Jadon, and Vikas Kumar. 2020 · 2019
Earlier work this paper cites.
Reducing gender bias in word-level language models with a gender-equalizing loss function
Yusu Qian, Urwa Muaz, Ben Zhang, and Jae Won Hyun. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Earlier work this paper cites.
What’s in a name? Reducing bias in bios without access to protected attributes
Alexey Romanov, Maria De-Arteaga, Hanna Wallach, Jennifer Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Geyik, Krishnaram Kenthapadi, Anna Rumshisky, and Adam Kalai. 2019 · 2019
Earlier work this paper cites.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Cynthia Rudin. 2019 · 2019
Earlier work this paper cites.
Gender bias in pretrained Swedish embeddings
Magnus Sahlgren and Fredrik Olsson. 2019 · 2019
Earlier work this paper cites.
The Risk of Racial Bias in Hate Speech Detection
Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A. Smith. 2019 · 2019
Earlier work this paper cites.
The role of protected class word lists in bias identification of contextualized word representations
João Sedoc and Lyle Ungar. 2019 · 2019
Earlier work this paper cites.
Evaluating Gender Bias in Machine Translation
Gabriel Stanovsky, Noah A. Smith, and Luke Zettlemoyer. 2019 · 2019
Earlier work this paper cites.
Mitigating Gender Bias in Natural Language Processing: Literature Review
Tony Sun, Andrew Gaut, Shirlyn Tang, Yuxin Huang, Mai ElSherief, Jieyu Zhao, Diba Mirza, Elizabeth Belding, Kai-Wei Chang, and William Yang Wang. 2019 · 2019
Earlier work this paper cites.
A transparent framework for evaluating unintended demographic bias in word embeddings
Chris Sweeney and Maryam Najafian. 2019 · 2019
Earlier work this paper cites.
Assessing Social and Intersectional Biases in Contextualized Word Representations
Yi Chern Tan and L. Elisa Celis. 2019 · 2019
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 2019
Earlier work this paper cites.
Gender bias in contextualized word embeddings
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Ryan Cotterell, Vicente Ordonez, and Kai-Wei Chang. 2019 · 2019
Earlier work this paper cites.
Mitigation of unintended biases against non-native english texts in sentiment analysis
Alina Zhiltsova, Simon Caton, and Catherine Mulway. 2019 · 2019
Earlier work this paper cites.
Examining Gender Bias in Languages with Grammatical Gender
Pei Zhou, Weijia Shi, Jieyu Zhao, Kuan-Hao Huang, Muhao Chen, Ryan Cotterell, and Kai-Wei Chang. 2019 · 2019
Earlier work this paper cites.
Counterfactual Data Augmentation for Mitigating Gender Stereotypes in Languages with Rich Morphology
Ran Zmigrod, Sebastian J. Mielke, Hanna Wallach, and Ryan Cotterell. 2019 · 2019
Earlier work this paper cites.
Unmasking Contextual Stereotypes: Measuring and Mitigating BERT’s Gender Bias
Marion Bartl, Malvina Nissim, and Albert Gatt. 2020 · 2020
Cited alongside, same era.
What is the point of fairness? disability, AI and the complexity of justice
Cynthia L. Bennett and Os Keyes. 2020 · 2020
Cited alongside, same era.
Language (technology) is power: a critical survey of “bias” in NLP
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020 · 2020
Cited alongside, same era.
Toward gender-inclusive coreference resolution
Yang Trista Cao and Hal Daumé III. 2020 · 2020
Cited alongside, same era.
Hindi-english hate speech detection: Author profiling, debiasing, and practical perspectives
Shivang Chopra, Ramit Sawhney, Puneet Mathur, and Rajiv Ratn Shah. 2020 · 2020
Cited alongside, same era.
A snapshot of the frontiers of fairness in machine learning
A Survey on Bias and Fairness in Natural Language Processing
Rajas Bansal. 2022 · 2022
Later among the works it cites.
A prompt array keeps the bias away: Debiasing vision-language models with adversarial learning
Hugo Berg, Siobhan Hall, Yash Bhalgat, Hannah Kirk, Aleksandar Shtedritski, and Max Bain. 2022 · 2022
Later among the works it cites.
Re-contextualizing fairness in NLP: The case of India
Shaily Bhatt, Sunipa Dev, Partha Talukdar, Shachi Dave, and Vinodkumar Prabhakaran. 2022 · 2022
Later among the works it cites.
On the intrinsic and extrinsic fairness evaluation metrics for contextualized language representations
Yang Cao, Yada Pruksachatkun, Kai-Wei Chang, Rahul Gupta, Varun Kumar, Jwala Dhamala, and Aram Galstyan. 2022 · 2022
Later among the works it cites.
Measuring fairness with biased rulers: A comparative study on bias metrics for pre-trained language models
Pieter Delobelle, Ewoenam Tokpo, Toon Calders, and Bettina Berendt. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alexandra Chouldechova and Aaron Roth. 2020 · 2020
Cited alongside, same era.
On measuring and mitigating biased inferences of word embeddings
Sunipa Dev, Tao Li, Jeff M. Phillips, and Vivek Srikumar. 2020 · 2020
Cited alongside, same era.
Unsupervised discovery of implicit gender bias
Anjalie Field and Yulia Tsvetkov. 2020 · 2020
Cited alongside, same era.
Measuring social bias in knowledge graph embeddings
Joseph Fisher, Dave Palfrey, Christos Christodoulopoulos, and Arpit Mittal. 2020 · 2020
Cited alongside, same era.
Evaluating models’ local decision boundaries via contrast sets
Matt Gardner, Yoav Artzi, Victoria Basmov, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, Nitish Gupta, Hannaneh Hajishirzi, Gabriel Ilharco, Daniel Khashabi, Kevin Lin, Jiangming Liu, Nelson F. Liu, Phoebe Mulcaire, Qiang Ning, Sameer Singh, Noah A. Smith, Sanjay Subramanian, Reut Tsarfaty, Eric Wallace, Ally Zhang, and Ben Zhou. 2020 · 2020
Cited alongside, same era.
Towards Understanding Gender Bias in Relation Extraction
Andrew Gaut, Tony Sun, Shirlyn Tang, Yuxin Huang, Jing Qian, Mai ElSherief, Jieyu Zhao, Diba Mirza, Elizabeth Belding, Kai-Wei Chang, and William Yang Wang. 2020 · 2020
Cited alongside, same era.
Towards a critical race methodology in algorithmic fairness
Alex Hanna, Emily Denton, Andrew Smart, and Jamila Smith-Loud. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
Theories of “gender” in nlp bias research
Hannah Devinney, Jenny Björklund, and Henrik Björklund. 2022 · 2022
Later among the works it cites.
Understanding gender bias in knowledge base embeddings
Yupei Du, Qi Zheng, Yuanbin Wu, Man Lan, Yan Yang, and Meirong Ma. 2022 · 2022
Later among the works it cites.
An ontology for fairness metrics
Jade S. Franklin, Karan Bhanot, Mohamed Ghalwash, Kristin P. Bennett, Jamie McCusker, and Deborah L. McGuinness. 2022 · 2022
Later among the works it cites.
Debiasing pretrained text encoders by paying attention to paying attention
Yacine Gaci, Boualem Benatallah, Fabio Casati, and Khalid Benabdeslem. 2022 · 2022
Later among the works it cites.
Kernel-whitening: Overcome dataset bias with isotropic sentence embedding
SongYang Gao, Shihan Dou, Qi Zhang, and Xuanjing Huang. 2022 · 2022
Later among the works it cites.
Auto-debias: Debiasing masked language models with automated biased prompts
Yue Guo, Yi Yang, and Ahmed Abbasi. 2022 · 2022
Later among the works it cites.
Mitigating gender bias in distilled language models via counterfactual role reversal
Umang Gupta, Jwala Dhamala, Varun Kumar, Apurv Verma, Yada Pruksachatkun, Satyapriya Krishna, Rahul Gupta, Kai-Wei Chang, Greg Ver Steeg, and Aram Galstyan. 2022a · 2022
Later among the works it cites.
Balancing out bias: Achieving fairness through balanced training
Xudong Han, Timothy Baldwin, and Trevor Cohn. 2022 · 2022
Later among the works it cites.
MABEL: Attenuating gender bias using textual entailment data
Jacqueline He, Mengzhou Xia, Christiane Fellbaum, and Danqi Chen. 2022a · 2022
Later among the works it cites.
Controlling bias exposure for fair interpretable predictions
Zexue He, Yu Wang, Julian McAuley, and Bodhisattwa Prasad Majumder. 2022b · 2022
Later among the works it cites.
Towards understanding gender-seniority compound bias in natural language generation
Samhita Honnavalli, Aesha Parekh, Lily Ou, Sophie Groenwold, Sharon Levy, Vicente Ordonez, and William Yang Wang. 2022 · 2022
Later among the works it cites.
Probing as quantifying inductive bias
Alexander Immer, Lucas Torroba Hennigen, Vincent Fortuin, and Ryan Cotterell. 2022 · 2022
Later among the works it cites.
Gender biases and where to find them: Exploring gender bias in pre-trained transformer-based language models using movement pruning
Przemyslaw Joniak and Akiko Aizawa. 2022 · 2022
Later among the works it cites.
Gender bias in meta-embeddings
Masahiro Kaneko, Danushka Bollegala, and Naoaki Okazaki. 2022a · 2022
Later among the works it cites.
Gender bias in masked language models for multiple languages
Masahiro Kaneko, Aizhan Imankulova, Danushka Bollegala, and Naoaki Okazaki. 2022b · 2022
Later among the works it cites.
Measuring fairness of text classifiers via prediction sensitivity
Satyapriya Krishna, Rahul Gupta, Apurv Verma, Jwala Dhamala, Yada Pruksachatkun, and Kai-Wei Chang. 2022 · 2022
Later among the works it cites.
An empirical study on pseudo-log-likelihood bias measures for masked language models using paraphrased sentences
Bum Chul Kwon and Nandana Mihindukulasooriya. 2022 · 2022
Later among the works it cites.
HERB: Measuring hierarchical regional bias in pre-trained language models
Yizhi Li, Ge Zhang, Bohao Yang, Chenghua Lin, Anton Ragni, Shi Wang, and Jie Fu. 2022 · 2022
Later among the works it cites.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher R’e, Diana Acosta-Navas, Drew A. Hudson, E. Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue Wang, Keshav Santhanam, Laurel J. Orr, Lucia Zheng, Mert Yuksekgonul, Mirac Suzgun, Nathan S. Kim, Neel Guha, Niladri S. Chatterji, Omar Khattab, Peter Henderson, Qian Huang, Ryan Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas F. Icard, Tianyi Zhang, Vishrav Chaudhary, William Wang, Xuechen Li, Yifan Mai, Yuhui Zhang, and Yuta Koreeda. 2022 · 2022
Later among the works it cites.
Don’t forget about pronouns: Removing gender bias in language models without losing factual gender information
Tomasz Limisiewicz and David Mareček. 2022 · 2022
Later among the works it cites.
Socially aware bias measurements for Hindi language representations
Vijit Malik, Sunipa Dev, Akihiro Nishi, Nanyun Peng, and Kai-Wei Chang. 2022 · 2022
Later among the works it cites.
Debiasing masks: A new framework for shortcut mitigation in NLU
Johannes Mario Meissner, Saku Sugawara, and Akiko Aizawa. 2022 · 2022
Later among the works it cites.
Pipelines for social bias testing of large language models
Debora Nozza, Federico Bianchi, and Dirk Hovy. 2022 · 2022
Later among the works it cites.
How gender debiasing affects internal model representations, and why it matters
Hadas Orgad, Seraphina Goldfarb-Tarrant, and Yonatan Belinkov. 2022 · 2022
Later among the works it cites.
Don’t just clean it, proxy clean it: Mitigating bias by proxy in pre-trained models
Swetasudha Panda, Ari Kobren, Michael Wick, and Qinlan Shen. 2022 · 2022
Later among the works it cites.
Counterfactually augmented data and unintended bias: The case of sexism and hate speech detection
Indira Sen, Mattia Samory, Claudia Wagner, and Isabelle Augenstein. 2022 · 2022
Later among the works it cites.
Quantifying social biases using templates is unreliable
Preethi Seshadri, Pouya Pezeshkpour, and Sameer Singh. 2022 · 2022
Later among the works it cites.
“I’m sorry to hear that”: Finding new biases in language models with a holistic descriptor dataset
Eric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani, and Adina Williams. 2022 · 2022
Later among the works it cites.
To prefer or to choose? generating agency and power counterfactuals jointly for gender bias mitigation
Maja Stahl, Maximilian Spliethöver, and Henning Wachsmuth. 2022 · 2022
Later among the works it cites.
Fewer errors, but more stereotypes? the effect of model size on gender bias
Yarden Tal, Inbal Magar, and Roy Schwartz. 2022 · 2022
Later among the works it cites.
A robust bias mitigation procedure based on the stereotype content model
Eddie Ungless, Amy Rafferty, Hrichika Nag, and Björn Ross. 2022 · 2022
Later among the works it cites.
The undesirable dependence on frequency of gender bias metrics based on word embeddings
Francisco Valentini, Germán Rosati, Diego Fernandez Slezak, and Edgar Altszyler. 2022 · 2022
Later among the works it cites.
A study of implicit bias in pretrained language models against people with disabilities
Pranav Narayanan Venkit, Mukund Srinath, and Shomir Wilson. 2022 · 2022
Later among the works it cites.
AutoCAD: Automatically generate counterfactuals for mitigating shortcut learning
Jiaxin Wen, Yeshuang Zhu, Jinchao Zhang, Jie Zhou, and Minlie Huang. 2022 · 2022
Later among the works it cites.
Interpreting the robustness of neural NLP models to textual perturbations
Yunxiang Zhang, Liangming Pan, Samson Tan, and Min-Yen Kan. 2022 · 2022
Later among the works it cites.
SODAPOP: Open-ended discovery of social biases in social commonsense reasoning models
Haozhe An, Zongxia Li, Jieyu Zhao, and Rachel Rudinger. 2023 · 2023
Closest in time.
A tale of pronouns: Interpretability informs gender bias mitigation for fairer instruction-tuned machine translation
Giuseppe Attanasio, Flor Plaza del Arco, Debora Nozza, and Anne Lauscher. 2023 · 2023
Closest in time.
Social commonsense for explanation and cultural bias discovery
Lisa Bauer, Hanna Tischer, and Mohit Bansal. 2023 · 2023
Closest in time.
Simplicity bias leads to amplified performance disparities
Samuel James Bell and Levent Sagun. 2023 · 2023
Closest in time.
How redundant are redundant encodings? blindness in the wild and racial disparity when race is unobserved
Lingwei Cheng, Isabel O Gallegos, Derek Ouyang, Jacob Goldin, and Dan Ho. 2023 · 2023
Closest in time.
Evaluation of African American language bias in natural language generation
Nicholas Deas, Jessica Grieser, Shana Kleiner, Desmond Patton, Elsbeth Turcan, and Kathleen McKeown. 2023 · 2023
Closest in time.
Detection and mitigation of algorithmic bias via predictive parity
Cyrus DiCiccio, Brian Hsu, Yinyin Yu, Preetam Nandy, and Kinjal Basu. 2023 · 2023
Closest in time.
Towards stable natural language understanding via information entropy guided debiasing
Li Du, Xiao Ding, Zhouhao Sun, Ting Liu, Bing Qin, and Jingshuo Liu. 2023 · 2023
Closest in time.
WinoQueer: A community-in-the-loop benchmark for anti-LGBTQ+ bias in large language models
Virginia Felkner, Ho-Chun Herbert Chang, Eugene Jang, and Jonathan May. 2023 · 2023
Closest in time.
Examining risks of racial biases in nlp tools for child protective services
Anjalie Field, Amanda Coston, Nupoor Gandhi, Alexandra Chouldechova, Emily Putnam-Hornstein, David Steier, and Yulia Tsvetkov. 2023 · 2023
Closest in time.
Cross-lingual transfer can worsen bias in sentiment analysis
Seraphina Goldfarb-Tarrant, Björn Ross, and Adam Lopez. 2023 · 2023
Closest in time.
Calm: A multi-task benchmark for comprehensive assessment of language model bias
Vipul Gupta, Pranav Narayanan Venkit, Hugo Laurençon, Shomir Wilson, and Rebecca J Passonneau. 2023 · 2023
Closest in time.
“fifty shades of bias”: Normative ratings of gender bias in GPT generated English text
Rishav Hada, Agrima Seth, Harshita Diddee, and Kalika Bali. 2023 · 2023
Closest in time.
Improving bias mitigation through bias experts in natural language understanding
Eojin Jeon, Mingyu Lee, Juhyeong Park, Yeachan Kim, Wing-Lam Mok, and SangKeun Lee. 2023 · 2023
Closest in time.
Comparing intrinsic gender bias evaluation measures without using human annotated examples
Masahiro Kaneko, Danushka Bollegala, and Naoaki Okazaki. 2023 · 2023
Closest in time.
Parameter-efficient modularised bias mitigation via AdapterFusion
Deepak Kumar, Oleg Lesota, George Zerveas, Daniel Cohen, Carsten Eickhoff, Markus Schedl, and Navid Rekabsaz. 2023 · 2023
Closest in time.
When do pre-training biases propagate to downstream tasks? a case study in text summarization
Faisal Ladhak, Esin Durmus, Mirac Suzgun, Tianyi Zhang, Dan Jurafsky, Kathleen McKeown, and Tatsunori Hashimoto. 2023 · 2023
Closest in time.
Target-agnostic gender-aware contrastive learning for mitigating bias in multilingual machine translation
Minwoo Lee, Hyukhun Koh, Kang-il Lee, Dongdong Zhang, Minsung Kim, and Kyomin Jung. 2023 · 2023
Closest in time.
Comparing biases and the impact of multilingual training across multiple languages
Sharon Levy, Neha John, Ling Liu, Yogarshi Vyas, Jie Ma, Yoshinari Fujinuma, Miguel Ballesteros, Vittorio Castelli, and Dan Roth. 2023 · 2023
Closest in time.
Prompt tuning pushes farther, contrastive learning pulls closer: A two-stage approach to mitigate social biases
Yingji Li, Mengnan Du, Xin Wang, and Ying Wang. 2023 · 2023
Closest in time.
Logic against bias: Textual entailment mitigates stereotypical sentence reasoning
Hongyin Luo and James Glass. 2023 · 2023
Closest in time.
Queenie Luo, Michael J Puett, and Michael D Smith. 2023 · 2023
Closest in time.
InterFair: Debiasing with natural language feedback for fair interpretable predictions
Bodhisattwa Majumder, Zexue He, and Julian McAuley. 2023 · 2023
Closest in time.
Bias against 93 stigmatized groups in masked language models and downstream sentiment classification tasks
Katelyn Mei, Sonia Fereidooni, and Aylin Caliskan. 2023 · 2023
Closest in time.
Socialstigmaqa: A benchmark to uncover stigma amplification in generative language models
Manish Nagireddy, Lamogha Chiazor, Moninder Singh, and Ioana Baldini. 2023 · 2023
Closest in time.
Towards a holistic approach: Understanding sociodemographic biases in nlp models using an interdisciplinary lens
Pranav Narayanan Venkit. 2023 · 2023
Closest in time.
Unmasking nationality bias: A study of human perception of nationalities in ai-generated articles
Pranav Narayanan Venkit, Sanjana Gautam, Ruchi Panchanadikar, Ting-Hao Huang, and Shomir Wilson. 2023 · 2023
Closest in time.
Social-group-agnostic bias mitigation via the stereotype content model
Ali Omrani, Alireza Salkhordeh Ziabari, Charles Yu, Preni Golazizian, Brendan Kennedy, Mohammad Atari, Heng Ji, and Morteza Dehghani. 2023 · 2023
Closest in time.
Evaluating biased attitude associations of language models in an intersectional context
Shiva Omrani Sabbaghi, Robert Wolfe, and Aylin Caliskan. 2023 · 2023
Closest in time.
BLIND: Bias removal with no demographics
Hadas Orgad and Yonatan Belinkov. 2023 · 2023
Closest in time.
The sins of the parents are to be laid upon the children: Biased humans, biased data, biased models
Merrick Osborne, Ali Omrani, and Morteza Dehghani. 2023 · 2023
Closest in time.
“i’m fully who i am”: Towards centering transgender and non-binary voices to measure biases in open language generation
Anaelia Ovalle, Palash Goyal, Jwala Dhamala, Zachary Jaggers, Kai-Wei Chang, Aram Galstyan, Richard Zemel, and Rahul Gupta. 2023 · 2023
Closest in time.
In-depth look at word filling societal bias measures
Matúš Pikuliak, Ivana Beňová, and Viktor Bachratý. 2023 · 2023
Closest in time.
Meta-analysis of the “ironic” effects of intergroup contact
Nils Karl Reimer and Nikhil Kumar Sengupta. 2023 · 2023
Closest in time.
Add-remove-or-relabel: Practitioner-friendly bias mitigation via influential fairness
Brianna Richardson, Prasanna Sattigeri, Dennis Wei, Karthikeyan Natesan Ramamurthy, Kush Varshney, Amit Dhurandhar, and Juan E. Gilbert. 2023 · 2023
Closest in time.
The tail wagging the dog: Dataset construction biases of social bias benchmarks
Nikil Selvam, Sunipa Dev, Daniel Khashabi, Tushar Khot, and Kai-Wei Chang. 2023 · 2023
Closest in time.
Representation bias in data: A survey on identification and resolution techniques
Nima Shahbazi, Yin Lin, Abolfazl Asudeh, and H. V. Jagadish. 2023 · 2023
Closest in time.
Language models get a gender makeover: Mitigating gender bias with few-shot data interventions
Himanshu Thakur, Atishay Jain, Praneetha Vaddamanu, Paul Pu Liang, and Louis-Philippe Morency. 2023 · 2023
Closest in time.
Diverse perspectives can mitigate political bias in crowdsourced content moderation
Jacob Thebault-Spieker, Sukrit Venkatagiri, Naomi Mine, and Kurt Luther. 2023 · 2023
Closest in time.
How far can it go? on intrinsic gender bias mitigation for text classification
Ewoenam Kwaku Tokpo, Pieter Delobelle, Bettina Berendt, and Toon Calders. 2023 · 2023
Closest in time.
Are fairy tales fair? analyzing gender bias in temporal narrative event chains of children’s fairy tales
Paulina Toro Isaza, Guangxuan Xu, Toye Oloko, Yufang Hou, Nanyun Peng, and Dakuo Wang. 2023 · 2023
Closest in time.
Measuring normative and descriptive biases in language models using census data
Samia Touileb, Lilja Øvrelid, and Erik Velldal. 2023 · 2023
Closest in time.
On the interpretability and significance of bias metrics in texts: a PMI-based approach
Francisco Valentini, Germán Rosati, Damián Blasi, Diego Fernandez Slezak, and Edgar Altszyler. 2023 · 2023
Closest in time.
The sentiment problem: A critical survey towards deconstructing sentiment analysis
Pranav Venkit, Mukund Srinath, Sanjana Gautam, Saranya Venkatraman, Vipul Gupta, Rebecca J. Passonneau, and Shomir Wilson. 2023a · 2023
Closest in time.
Counter-gap: Counterfactual bias evaluation through gendered ambiguous pronouns
Zhongbin Xie, Vid Kocijan, Thomas Lukasiewicz, and Oana-Maria Camburu. 2023 · 2023
Closest in time.
Conceptor-aided debiasing of large language models
Li Yifei, Lyle Ungar, and João Sedoc. 2023 · 2023
Closest in time.
Model debiasing via gradient-based explanation on representation
Jindi Zhang, Luning Wang, Dan Su, Yongxiang Huang, Caleb Chen Cao, and Lei Chen. 2023 · 2023
Closest in time.
A predictive factor analysis of social biases and task-performance in pretrained masked language models
Yi Zhou, Jose Camacho-Collados, and Danushka Bollegala. 2023b · 2023
Closest in time.
Exploring ai ethics of chatgpt: A diagnostic analysis
Terry Yue Zhuo, Yujin Huang, Chunyang Chen, and Zhenchang Xing. 2023 · 2023
Closest in time.
The (undesired) attenuation of human biases by multilinguality
Cristina España-Bonet and Alberto Barrón-Cedeño. 2022 · 2077
Closest in time.
BBQ: A hand-built bias benchmark for question answering
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel Bowman. 2022a · 2086
Closest in time.
BBQ: A hand-built bias benchmark for question answering
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel Bowman. 2022b · 2086
Closest in time.
No word embedding model is perfect: Evaluating the representation accuracy for social bias in the media
Maximilian Spliethöver, Maximilian Keiff, and Henning Wachsmuth. 2022 · 2093
Closest in time.