Fetching the paper…
Reading the bibliography…
Fine-tuned language models have been shown to exhibit biases against protected groups in a host of modeling tasks such as text classification and coreference resolution.
Transfer of machine learning fairness across domains
Candice Schumann, Xuezhi Wang, Alex Beutel, J. Chen, Hai Qian, and Ed Huai hsin Chi. 2019 · 1906
Earlier work this paper cites.
Investigating gender bias in bert
Rishabh Bhardwaj, Navonil Majumder, and Soujanya Poria. 2020 · 2009
Earlier work this paper cites.
A survey on transfer learning
Sinno Jialin Pan and Qiang Yang. 2010 · 2010
Earlier work this paper cites.
Ontonotes release 5.0 ldc2013t19
Ralph Weischedel, Martha Palmer, Mitchell Marcus, Eduard Hovy, Sameer Pradhan, Lance Ramshaw, Nianwen Xue, Ann Taylor, Jeff Kaufman, Michelle Franchini, et al. 2013 · 2013
Earlier work this paper cites.
Semi-supervised sequence learning
Andrew M. Dai and Quoc V. Le. 2015 · 2015
Earlier work this paper cites.
Demographic dialectal variation in social media: A case study of African-American English
Su Lin Blodgett, Lisa Green, and Brendan O’Connor. 2016 · 2016
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y. Zou, Venkatesh Saligrama, and Adam Tauman Kalai. 2016 · 2016
Earlier work this paper cites.
Equality of opportunity in supervised learning
Moritz Hardt, Eric Price, and Nati Srebro. 2016 · 2016
Earlier work this paper cites.
Data decisions and theoretical implications when adversarially learning fair representations
Alex Beutel, J. Chen, Zhe Zhao, and Ed Huai hsin Chi. 2017 · 2017
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
A. Caliskan, J. Bryson, and A. Narayanan. 2017 · 2017
Earlier work this paper cites.
Automated hate speech detection and the problem of offensive language
Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017 · 2017
Earlier work this paper cites.
Hate speech dataset from a white supremacy forum
Ona de Gibert, Naiara Perez, Aitor García-Pablos, and Montse Cuadros. 2018 · 2018
Earlier work this paper cites.
Measuring and mitigating unintended bias in text classification
Lucas Dixon, John Li, Jeffrey Scott Sorensen, Nithum Thain, and L. Vasserman. 2018 · 2018
Earlier work this paper cites.
Adversarial removal of demographic attributes from text data
Yanai Elazar and Yoav Goldberg. 2018 · 2018
Cited alongside, same era.
Large scale crowdsourcing and characterization of twitter abusive behavior
Antigoni-Maria Founta, Constantinos Djouvas, Despoina Chatzakou, I. Leontiadis, Jeremy Blackburn, G. Stringhini, Athena Vakali, M. Sirivianos, and Nicolas Kourtellis. 2018 · 2018
Cited alongside, same era.
The gab hate corpus: A collection of 27k posts annotated for hate speech
Brendan Kennedy, Mohammad Atari, Aida M Davani, Leigh Yeh, Ali Omrani, Yehsong Kim, Kris Coombs Jr, Shreya Havaldar, Gwenyth Portillo-Wightman, Elaine Gonzalez, et al. 2018 · 2018
Cited alongside, same era.
Explicit inductive bias for transfer learning with convolutional networks
Xuhong Li, Yves Grandvalet, and Franck Davoine. 2018 · 2018
Cited alongside, same era.
Learning adversarially fair and transferable representations
David Madras, Elliot Creager, Toniann Pitassi, and Richard S. Zemel. 2018 · 2018
Cited alongside, same era.
On measuring social biases in sentence encoders
Chandler May, Alex Wang, Shikha Bordia, Samuel R. Bowman, and Rachel Rudinger. 2019 · 2019
Later among the works it cites.
Debiasing embeddings for reduced gender bias in text classification
Flavien Prost, Nithum Thain, and Tolga Bolukbasi. 2019 · 2019
Later among the works it cites.
The risk of racial bias in hate speech detection
Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A. Smith. 2019 · 2019
Later among the works it cites.
Gender bias in contextualized word embeddings
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Ryan Cotterell, Vicente Ordonez, and Kai-Wei Chang. 2019 · 2019
Later among the works it cites.
Examining gender bias in languages with grammatical gender
Pei Zhou, Weijia Shi, Jieyu Zhao, Kuan-Hao Huang, Muhao Chen, Ryan Cotterell, and Kai-Wei Chang. 2019 · 2019
Later among the works it cites.
Language (technology) is power: A critical survey of “bias” in NLP
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reducing gender bias in abusive language detection
Ji Ho Park, Jamin Shin, and Pascale Fung. 2018 · 2018
Cited alongside, same era.
Mitigating unwanted biases with adversarial learning
B. H. Zhang, B. Lemoine, and Margaret Mitchell. 2018 · 2018
Cited alongside, same era.
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018 · 2018
Cited alongside, same era.
Bias in bios: A case study of semantic representation bias in a high-stakes setting
Maria De-Arteaga, Alexey Romanov, Hanna Wallach, Jennifer Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Geyik, Krishnaram Kenthapadi, and Adam Tauman Kalai. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
BERT for coreference resolution: Baselines and analysis
Mandar Joshi, Omer Levy, Luke Zettlemoyer, and Daniel Weld. 2019 · 2019
Cited alongside, same era.
Topics to avoid: Demoting latent confounds in text classification
Sachin Kumar, Shuly Wintner, Noah A. Smith, and Yulia Tsvetkov. 2019 · 2019
Cited alongside, same era.
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020 · 2020
Closest in time.
Contextualizing hate speech classifiers with post-hoc explanation
Brendan Kennedy, Xisen Jin, Aida Mostafazadeh Davani, Morteza Dehghani, and Xiang Ren. 2020 · 2020
Closest in time.
Towards debiasing sentence representations
Paul Pu Liang, Irene Mengze Li, Emily Zheng, Yao Chong Lim, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2020 · 2020
Closest in time.
Null it out: Guarding protected attributes by iterative nullspace projection
Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg. 2020 · 2020
Closest in time.
Adversarially robust transfer learning
Ali Shafahi, Parsa Saadatpanah, Chen Zhu, Amin Ghiasi, Christoph Studer, David W. Jacobs, and Tom Goldstein. 2020 · 2020
Closest in time.
Demoting racial bias in hate speech detection
Mengzhou Xia, Anjalie Field, and Yulia Tsvetkov. 2020 · 2020
Closest in time.
Demographics should not be the reason of toxicity: Mitigating discrimination in text classifications with instance weighting
Guanhua Zhang, Bing Bai, Junqi Zhang, Kun Bai, Conghui Zhu, and Tiejun Zhao. 2020 · 2020
Closest in time.