Language models are few-shot learners
Brown, Tom, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
RoBERTa: A robustly optimized bert pretraining approach
Original
Liu, Yinhan, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Plug and play language models: A simple approach to controlled text generation
Original
Dathathri, Sumanth, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. 2019 · 1912
Earlier work this paper cites.
Intrinsic bias metrics do not correlate with application bias
Goldfarb-Tarrant, Seraphina, Rebecca Marchant, Ricardo Muñoz Sánchez, Mugdha Pandya, and Adam Lopez. 2021 · 1940
Earlier work this paper cites.
Multicalibration: Calibration for the (computationally-identifiable) masses
Hébert-Johnson, Ursula, Michael Kim, Omer Reingold, and Guy Rothblum. 2018 · 1948
Earlier work this paper cites.
RedditBias: A real-world resource for bias evaluation and debiasing of conversational language models
Barikeri, Soumya, Anne Lauscher, Ivan Vulić, and Goran Glavaš. 2021 · 1955
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, Nitish, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
Measuring individual differences in implicit cognition: The implicit association test
Greenwald, Anthony G, Debbie E McGhee, and Jordan LK Schwartz. 1998 · 1998
Earlier work this paper cites.
Linguistic intergroup bias: Stereotype perpetuation through language
Maass, Anne. 1999 · 1999
Earlier work this paper cites.
Racial identification by speech
Baugh, John. 2000 · 2000
Earlier work this paper cites.
Bringing the people back in: Contesting benchmark machine learning datasets
Original
Denton, Emily, Alex Hanna, Razvan Amironesei, Andrew Smart, Hilary Nicole, and Morgan Klaus Scheuerman. 2020 · 2007
Earlier work this paper cites.
Measuring and reducing gendered correlations in pre-trained models
Original
Webster, Kellie, Xuezhi Wang, Ian Tenney, Alex Beutel, Emily Pitler, Ellie Pavlick, Jilin Chen, Ed Chi, and Slav Petrov. 2020 · 2010
Earlier work this paper cites.
Recipes for safety in open-domain chatbots
Original
Xu, Jing, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan. 2020 · 2010
Earlier work this paper cites.
Fairness through awareness
Dwork, Cynthia, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Richard Zemel. 2012 · 2012
Earlier work this paper cites.
Moral responsibility for computing artifacts: "The rules" and issues of trust
Grodzinsky, F. S., K. Miller, and M. J. Wolf. 2012 · 2012
Earlier work this paper cites.
Data preprocessing techniques for classification without discrimination
Kamiran, Faisal and Toon Calders. 2012 · 2012
Earlier work this paper cites.
The Winograd schema challenge
Levesque, Hector, Ernest Davis, and Leora Morgenstern. 2012 · 2012
Earlier work this paper cites.
Implicit attitudes and the perception of sociolinguistic variation
Loudermilk, Brandon C. 2015 · 2015
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? Debiasing word embeddings
Bolukbasi, Tolga, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. 2016 · 2016
Earlier work this paper cites.
Equality of opportunity in supervised learning
Hardt, Moritz, Eric Price, and Nati Srebro. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Rajpurkar, Pranav, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
"Why should I trust you?" Explaining the predictions of any classifier
Ribeiro, Marco Tulio, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Racial disparity in natural language processing: A case study of social media African-American English
Original
Blodgett, Su Lin and Brendan O’Connor. 2017 · 2017
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Caliskan, Aylin, Joanna J. Bryson, and Arvind Narayanan. 2017 · 2017
Earlier work this paper cites.
SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Cer, Daniel, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Earlier work this paper cites.
Fair prediction with disparate impact: A study of bias in recidivism prediction instruments
Chouldechova, Alexandra. 2017 · 2017
Earlier work this paper cites.
The trouble with bias
Crawford, Kate. 2017 · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, James, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. 2017 · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Lundberg, Scott M and Su-In Lee. 2017 · 2017
Earlier work this paper cites.
Men also like shopping: Reducing gender bias amplification using corpus-level constraints
Zhao, Jieyu, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2017 · 2017
Earlier work this paper cites.
Hurtlex: A multilingual lexicon of words to hurt
Bassignana, Elisa, Valerio Basile, Viviana Patti, et al. 2018 · 2018
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Bender, Emily M and Batya Friedman. 2018 · 2018
Earlier work this paper cites.
Measuring and mitigating unintended bias in text classification
Dixon, Lucas, John Li, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2018 · 2018
Earlier work this paper cites.
Word embeddings quantify 100 years of gender and ethnic stereotypes
Garg, Nikhil, Londa Schiebinger, Dan Jurafsky, and James Zou. 2018 · 2018
Earlier work this paper cites.
Preventing fairness gerrymandering: Auditing and learning for subgroup fairness
Kearns, Michael, Seth Neel, Aaron Roth, and Zhiwei Steven Wu. 2018 · 2018
Earlier work this paper cites.
Examining gender and race bias in two hundred sentiment analysis systems
Kiritchenko, Svetlana and Saif Mohammad. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, Alec, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. 2018 · 2018
Earlier work this paper cites.
Gender bias in coreference resolution
Rudinger, Rachel, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. 2018 · 2018
Earlier work this paper cites.
Fairness definitions explained
Verma, Sahil and Julia Rubin. 2018 · 2018
Earlier work this paper cites.
Mind the GAP: A balanced corpus of gendered ambiguous pronouns
Webster, Kellie, Marta Recasens, Vera Axelrod, and Jason Baldridge. 2018 · 2018
Earlier work this paper cites.
Mitigating unwanted biases with adversarial learning
Zhang, Brian Hu, Blake Lemoine, and Margaret Mitchell. 2018 · 2018
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Zhang, Hongyi, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz. 2018 · 2018
Earlier work this paper cites.
Gender bias in coreference resolution: Evaluation and debiasing methods
Zhao, Jieyu, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018 · 2018
Earlier work this paper cites.
Fairness and Machine Learning: Limitations and Opportunities
Barocas, Solon, Moritz Hardt, and Arvind Narayanan. 2019 · 2019
Earlier work this paper cites.
A typology of ethical risks in language technology with an eye towards where transparent documentation can help
Bender, Emily M. 2019 · 2019
Earlier work this paper cites.
How stereotypes are shared through language: a review and introduction of the social categories and stereotypes communication (SCSC) framework
Beukeboom, Camiel J and Christian Burgers. 2019 · 2019
Earlier work this paper cites.
Identifying and reducing gender bias in word-level language models
Bordia, Shikha and Samuel R. Bowman. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, Jacob, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Counterfactual fairness in text classification through robustness
Garg, Sahaj, Vincent Perot, Nicole Limtiaco, Ankur Taly, Ed H. Chi, and Alex Beutel. 2019 · 2019
Earlier work this paper cites.
Lipstick on a pig: Debiasing methods cover up systematic gender biases in word embeddings but do not remove them
Gonen, Hila and Yoav Goldberg. 2019 · 2019
Earlier work this paper cites.
"Good" isn’t good enough
Green, Ben. 2019 · 2019
Earlier work this paper cites.
It’s all in the name: Mitigating gender bias with name-based counterfactual data substitution
Hall Maudslay, Rowan, Hila Gonen, Ryan Cotterell, and Simone Teufel. 2019 · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
Houlsby, Neil, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Earlier work this paper cites.
Measuring bias in contextualized word representations
Kurita, Keita, Nidhi Vyas, Ayush Pareek, Alan W Black, and Yulia Tsvetkov. 2019 · 2019
Earlier work this paper cites.
Black is to criminal as Caucasian is to police: Detecting and removing multiclass bias in word embeddings
Manzini, Thomas, Lim Yao Chong, Alan W Black, and Yulia Tsvetkov. 2019 · 2019
Earlier work this paper cites.
On measuring social biases in sentence encoders
May, Chandler, Alex Wang, Shikha Bordia, Samuel R. Bowman, and Rachel Rudinger. 2019 · 2019
Earlier work this paper cites.
Model cards for model reporting
Mitchell, Margaret, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019 · 2019
Earlier work this paper cites.
Reducing gender bias in word-level language models with a gender-equalizing loss function
Qian, Yusu, Urwa Muaz, Ben Zhang, and Jae Won Hyun. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, Alec, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Earlier work this paper cites.
The risk of racial bias in hate speech detection
Sap, Maarten, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A. Smith. 2019 · 2019
Earlier work this paper cites.
The woman worked as a babysitter: On biases in language generation
Sheng, Emily, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. 2019 · 2019
Earlier work this paper cites.
Assessing social and intersectional biases in contextualized word representations
Tan, Yi Chern and L. Elisa Celis. 2019 · 2019
Earlier work this paper cites.
Indigenous data, indigenous methodologies and indigenous data sovereignty
Walter, Maggie and Michele Suina. 2019 · 2019
Earlier work this paper cites.
BERT has a mouth, and it must speak: BERT as a Markov random field language model
Wang, Alex and Kyunghyun Cho. 2019 · 2019
Earlier work this paper cites.
Gender bias in contextualized word embeddings
Zhao, Jieyu, Tianlu Wang, Mark Yatskar, Ryan Cotterell, Vicente Ordonez, and Kai-Wei Chang. 2019 · 2019
Earlier work this paper cites.
Counterfactual data augmentation for mitigating gender stereotypes in languages with rich morphology
Zmigrod, Ran, Sabrina J. Mielke, Hanna Wallach, and Ryan Cotterell. 2019 · 2019
Earlier work this paper cites.
Unmasking contextual stereotypes: Measuring and mitigating BERT’s gender bias
Bartl, Marion, Malvina Nissim, and Albert Gatt. 2020 · 2020
Earlier work this paper cites.
Race After Technology: Abolitionist Tools for the New Jim Code
Benjamin, Ruha. 2020 · 2020
Earlier work this paper cites.
Language (technology) is power: A critical survey of “bias” in NLP
Blodgett, Su Lin, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020 · 2020
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Conneau, Alexis, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Earlier work this paper cites.
Language and discrimination: Generating meaning, perceiving identities, and discriminating outcomes
Craft, Justin T, Kelly E Wright, Rachel Elizabeth Weissler, and Robin M Queen. 2020 · 2020
Earlier work this paper cites.
Detecting gender stereotypes: Lexicon vs. supervised learning methods
Cryan, Jenna, Shiliang Tang, Xinyi Zhang, Miriam Metzger, Haitao Zheng, and Ben Y Zhao. 2020 · 2020
Earlier work this paper cites.
On measuring and mitigating biased inferences of word embeddings
Dev, Sunipa, Tao Li, Jeff M Phillips, and Vivek Srikumar. 2020 · 2020
Earlier work this paper cites.
Queens are powerful too: Mitigating gender bias in dialogue generation
Dinan, Emily, Angela Fan, Adina Williams, Jack Urbanek, Douwe Kiela, and Jason Weston. 2020 · 2020
Earlier work this paper cites.
Social chemistry 101: Learning to reason about social and moral norms
Forbes, Maxwell, Jena D. Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi. 2020 · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Gehman, Samuel, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020 · 2020
Earlier work this paper cites.
Towards a critical race methodology in algorithmic fairness
Hanna, Alex, Emily Denton, Andrew Smart, and Jamila Smith-Loud. 2020 · 2020
Earlier work this paper cites.
Reducing sentiment bias in language models via counterfactual evaluation
Huang, Po-Sen, Huan Zhang, Ray Jiang, Robert Stanforth, Johannes Welbl, Jack Rae, Vishal Maini, Dani Yogatama, and Pushmeet Kohli. 2020 · 2020
Earlier work this paper cites.
Social biases in NLP models as barriers for persons with disabilities
Hutchinson, Ben, Vinodkumar Prabhakaran, Emily Denton, Kellie Webster, Yu Zhong, and Stephen Denuyl. 2020 · 2020
Earlier work this paper cites.
Mitigating gender bias amplification in distribution by posterior regularization
Jia, Shengyu, Tao Meng, Jieyu Zhao, and Kai-Wei Chang. 2020 · 2020
Earlier work this paper cites.
Don’t ask if artificial intelligence is good or fair, ask how it shifts power
Kalluri, Pratyusha et al. 2020 · 2020
Earlier work this paper cites.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Lewis, Mike, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Earlier work this paper cites.
UNQOVERing stereotyping biases via underspecified questions
Li, Tao, Daniel Khashabi, Tushar Khot, Ashish Sabharwal, and Vivek Srikumar. 2020 · 2020
Earlier work this paper cites.
Towards debiasing sentence representations
Liang, Paul Pu, Irene Mengze Li, Emily Zheng, Yao Chong Lim, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2020 · 2020
Earlier work this paper cites.
Does gender matter? Towards fairness in dialogue systems
Liu, Haochen, Jamell Dacon, Wenqi Fan, Hui Liu, Zitao Liu, and Jiliang Tang. 2020 · 2020
Earlier work this paper cites.
Gender bias in neural natural language processing
Lu, Kaiji, Piotr Mardziel, Fangjing Wu, Preetam Amancharla, and Anupam Datta. 2020 · 2020
Earlier work this paper cites.
PowerTransformer: Unsupervised controllable revision for biased language correction
Ma, Xinyao, Maarten Sap, Hannah Rashkin, and Yejin Choi. 2020 · 2020
Earlier work this paper cites.
Hate speech detection and racial bias mitigation in social media based on bert model
Mozafari, Marzieh, Reza Farahbakhsh, and Noël Crespi. 2020 · 2020
Earlier work this paper cites.