Fetching the paper…
Reading the bibliography…
Pretrained neural language models (LMs) are prone to generating racist, sexist, or otherwise toxic language which hinders their safe deployment.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner. 1998 · 1998
Earlier work this paper cites.
African American English: A Linguistic Introduction , 8.3.2002 edition edition
Lisa Green. 2002 · 2002
Earlier work this paper cites.
From user-centered to participatory design approaches , pages 1–7
Elizabeth Sanders. 2002 · 2002
Earlier work this paper cites.
Transforming classifier scores into accurate multiclass probability estimates
Bianca Zadrozny and Charles Elkan. 2002 · 2002
Earlier work this paper cites.
Empathy, ways of knowing, and interdependence as mediators of gender differences in attitudes toward hate speech and freedom of speech
Gloria Cowan and Désirée Khatchadourian. 2003 · 2003
Earlier work this paper cites.
Value sensitive design and information systems
Batya Friedman, Peter H Kahn, and Alan Borning. 2008 · 2008
Earlier work this paper cites.
Sparse additive generative models of text
Jacob Eisenstein, Amr Ahmed, and Eric P. Xing. 2011 · 2011
Earlier work this paper cites.
Mining of massive datasets
Anand Rajaraman and Jeffrey David Ullman. 2011 · 2011
Earlier work this paper cites.
Communities: Participatory design for, with and by communities
Carl DiSalvo, Andrew Clement, and Volkmar Pipek. 2012 · 2012
Earlier work this paper cites.
A computational approach to politeness with application to social factors
Cristian Danescu-Niculescu-Mizil, Moritz Sudhof, Dan Jurafsky, Jure Leskovec, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Earlier work this paper cites.
Demographic dialectal variation in social media: A case study of African-American English
Su Lin Blodgett, Lisa Green, and Brendan O’Connor. 2016 · 2016
Earlier work this paper cites.
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2016 · 2016
Earlier work this paper cites.
Ethan Fast, Tina Vachovsky, and Michael S. Bernstein. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Are you a racist or am I seeing things? annotator influence on hate speech detection on Twitter
Zeerak Waseem. 2016 · 2016
Earlier work this paper cites.
The problem with bias: Allocative versus representational harms in machine learning
Solon Barocas, Kate Crawford, Aaron Shapiro, and Hanna Wallach. 2017 · 2017
Earlier work this paper cites.
Controlling linguistic style aspects in neural language generation
Jessica Ficler and Yoav Goldberg. 2017 · 2017
Earlier work this paper cites.
Affect-LM: A neural language model for customizable affective text generation
Sayan Ghosh, Mathieu Chollet, Eugene Laksana, Louis-Philippe Morency, and Stefan Scherer. 2017 · 2017
Earlier work this paper cites.
A large labeled corpus for online harassment research
Jennifer Golbeck, Zahra Ashktorab, Rashad O. Banjo, Alexandra Berlinger, Siddharth Bhagwan, Cody Buntain, Paul Cheakalos, Alicia A. Geller, Quint Gergory, Rajesh Kumar Gnanasekaran, Raja Rajan Gunasekaran, Kelly M. Hoffman, Jenny Hottle, Vichita Jienjitlert, Shivika Khare, Ryan Lau, Marianna J. Martindale, Shalmali Naik, Heather L. Nixon, Piyush Ramachandran, Kristine M. Rogers, Lisa Rogers, Meghna Sardana Sarin, Gaurav Shahane, Jayanee Thanki, Priyanka Vengataraman, Zijian Wan, and Derek Michael Wu. 2017 · 2017
Earlier work this paper cites.
spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing
Matthew Honnibal and Ines Montani. 2017 · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell. 2017 · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang. 2017 · 2017
Earlier work this paper cites.
#gamergate and the fappening: How Reddit’s algorithm, governance, and culture support toxic technocultures
Adrienne Massanari. 2017 · 2017
Earlier work this paper cites.
The impact of toxic language on the health of Reddit communities
Shruthi Mohan, Apala Guha, Michael Harris, Fred Popowich, Ashley Schuster, and Chris Priebe. 2017 · 2017
Earlier work this paper cites.
One-step and two-step classification for abusive language detection on Twitter
Ji Ho Park and Pascale Fung. 2017 · 2017
Earlier work this paper cites.
Measuring the reliability of hate speech annotations: the case of the european refugee crisis
Björn Ross, Michael Rist, Guillermo Carbonell, Benjamin Cabrera, Nils Kurowsky, and Michael Wojatzki. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, undefinedukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Predicting factuality of reporting and bias of news media sources
Ramy Baly, Georgi Karadzhov, Dimitar Alexandrov, James Glass, and Preslav Nakov. 2018 · 2018
Cited alongside, same era.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M. Bender and Batya Friedman. 2018 · 2018
Cited alongside, same era.
Measuring and mitigating unintended bias in text classification
Lucas Dixon, John Li, Jeffrey Scott Sorensen, Nithum Thain, and Lucy Vasserman. 2018 · 2018
Cited alongside, same era.
Large scale crowdsourcing and characterization of Twitter abusive behavior
Antigoni-Maria Founta, Constantinos Djouvas, Despoina Chatzakou, Ilias Leontiadis, Jeremy Blackburn, Gianluca Stringhini, Athena Vakali, Michael Sirivianos, and Nicolas Kourtellis. 2018 · 2018
Cited alongside, same era.
Learning to write with cooperative discriminators
Ari Holtzman, Jan Buys, Maxwell Forbes, Antoine Bosselut, David Golub, and Yejin Choi. 2018 · 2018
Cited alongside, same era.
The woman worked as a babysitter: On biases in language generation
Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. 2019 · 2019
Later among the works it cites.
Detecting harassment in real-time as conversations develop
Wessel Stoop, Florian Kunneman, Antal van den Bosch, and Ben Miller. 2019 · 2019
Later among the works it cites.
“transforming” delete, retrieve, generate approach for controlled text style transfer
Akhilesh Sudhakar, Bhargav Upadhyay, and Arjun Maheswaran. 2019 · 2019
Later among the works it cites.
Universal adversarial triggers for attacking and analyzing NLP
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2019 · 2019
Later among the works it cites.
Bert has a mouth, and it must speak: Bert as a markov random field language model
Alex Wang and Kyunghyun Cho. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Cited alongside, same era.
Conversations gone awry: Detecting early signs of conversational failure
Justine Zhang, Jonathan Chang, Cristian Danescu-Niculescu-Mizil, Lucas Dixon, Yiqing Hua, Dario Taraborelli, and Nithum Thain. 2018 · 2018
Cited alongside, same era.
Evaluating the underlying gender bias in contextualized word embeddings
Christine Basta, Marta R. Costa-jussà, and Noe Casas. 2019 · 2019
Cited alongside, same era.
Finding microaggressions in the wild: A case for locating elusive phenomena in social media posts
Luke Breitfeller, Emily Ahn, David Jurgens, and Yulia Tsvetkov. 2019 · 2019
Cited alongside, same era.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Xiaodong Song. 2019 · 2019
Cited alongside, same era.
Gmail smart compose: Real-time assisted writing
Mia Xu Chen, Benjamin N. Lee, Gagan Bansal, Yuan Cao, Shuyuan Zhang, Justin Y. Lu, Jackie Tsay, Yinan Wang, Andrew M. Dai, Zhifeng Chen, Timothy Sohn, and Yonghui Wu. 2019 · 2019
Cited alongside, same era.
How automated tools discriminate against black language
Anna Chung. 2019 · 2019
Cited alongside, same era.
Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, and Jason Weston. 2019 · 2019
Later among the works it cites.
HuggingFace’s Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew. 2019 · 2019
Later among the works it cites.
Defending against neural fake news
Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. 2019 · 2019
Later among the works it cites.
Gender bias in contextualized word embeddings
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Ryan Cotterell, Vicente Ordonez, and Kai-Wei Chang. 2019 · 2019
Later among the works it cites.
Discovering and categorising language biases in Reddit
Xavier Ferrer Aran, T. V. Nuenen, J. M. Such, and N. Criado. 2020 · 2020
Closest in time.
Seven-in-Ten Reddit users get news on the site
Michael Barthel, Galen Stocking, Jesse Holcomb, and Amy Mitchell. 2016 · 2020
Closest in time.
Language (technology) is power: A critical survey of “bias” in NLP
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020 · 2020
Closest in time.
Language models are Few-Shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Closest in time.
Plug and play language models: A simple approach to controlled text generation
Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. 2020 · 2020
Closest in time.
Bringing the people back in: Contesting benchmark machine learning datasets
Emily Denton, Alex Hanna, Razvan Amironesei, Andrew Smart, Hilary Nicole, and Morgan Klaus Scheuerman. 2020 · 2020
Closest in time.
Multi-dimensional gender bias classification
Emily Dinan, A. Fan, Ledell Yu Wu, J. Weston, Douwe Kiela, and Adina Williams. 2020 · 2020
Closest in time.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020 · 2020
Closest in time.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Closest in time.
Social biases in NLP models as barriers for persons with disabilities
Ben Hutchinson, Vinodkumar Prabhakaran, Emily Denton, Kellie Webster, Yu Zhong, and Stephen Denuyl. 2020 · 2020
Closest in time.
Lessons from archives: Strategies for collecting sociocultural data in machine learning
Eun Seo Jo and Timnit Gebru. 2020 · 2020
Closest in time.
Talk to Transformer
Adam King. 2019 · 2020
Closest in time.
PowerTransformer: Unsupervised controllable revision for biased language correction
Xinyao Ma, Maarten Sap, Hannah Rashkin, and Yejin Choi. 2020 · 2020
Closest in time.
The radicalization risks of GPT-3 and advanced neural language models
Kris McGuffie and Alex Newhouse. 2020 · 2020
Closest in time.
Quick, community-specific learning: How distinctive toxicity norms are maintained in political subreddits
Ashwin Rajadesingan, Paul Resnick, and Ceren Budak. 2020 · 2020
Closest in time.
Reddit just banned one of its most toxic forums. but it won’t touch The_Donald
Aja Romano. 2017 · 2020
Closest in time.
Social bias frames: Reasoning about social and power implications of language
Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. 2020 · 2020
Closest in time.
Know thy corpus! robust methods for digital curation of web corpora
Serge Sharoff. 2020 · 2020
Closest in time.