Fetching the paper…
Reading the bibliography…
Biased associations have been a challenge in the development of classifiers for detecting toxic language, hindering both fairness and accuracy.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Linguistic politeness: current research issues
Gabriele Kasper. 1990 · 1990
Earlier work this paper cites.
African-American language use: Ideology and so-called obscenity
Arthur K Spears. 1998 · 1998
Earlier work this paper cites.
African American English: A Linguistic Introduction
Lisa Green. 2002 · 2002
Earlier work this paper cites.
Swearing methodologically : the (im)politeness of expletives in anonymous commentaries on youtube
Marta Dynel. 2012 · 2012
Earlier work this paper cites.
How to do things with slurs: Studies in the way of derogatory words
Adam M Croom. 2013 · 2013
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
The landscape of impoliteness research
Marta Dynel. 2015 · 2015
Earlier work this paper cites.
Demographic dialectal variation in social media: A case study of African-American English
Su Lin Blodgett, Lisa Green, and Brendan O’Connor. 2016 · 2016
Earlier work this paper cites.
Equality of opportunity in supervised learning
Moritz Hardt, Eric Price, and Nati Srebro. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Automated hate speech detection and the problem of offensive language
Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017 · 2017
Earlier work this paper cites.
Unsettling race and language: Toward a raciolinguistic perspective
Jonathan Rosa and Nelson Flores. 2017 · 2017
Earlier work this paper cites.
Measuring the reliability of hate speech annotations: the case of the european refugee crisis
Björn Ross, Michael Rist, Guillermo Carbonell, Benjamin Cabrera, Nils Kurowsky, and Michael Wojatzki. 2017 · 2017
Earlier work this paper cites.
The effect of different writing tasks on linguistic style: A case study of the roc story cloze task
Roy Schwartz, Maarten Sap, Ioannis Konstas, Li Zilles, Yejin Choi, and Noah A Smith. 2017 · 2017
Earlier work this paper cites.
Measuring and mitigating unintended bias in text classification
Lucas Dixon, John Li, Jeffrey Scott Sorensen, Nithum Thain, and L. Vasserman. 2018 · 2018
Earlier work this paper cites.
Large scale crowdsourcing and characterization of twitter abusive behavior
Antigoni-Maria Founta, Constantinos Djouvas, Despoina Chatzakou, Ilias Leontiadis, Jeremy Blackburn, Gianluca Stringhini, Athena Vakali, Michael Sirivianos, and Nicolas Kourtellis. 2018 · 2018
Earlier work this paper cites.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel R. Bowman, and Noah A. Smith. 2018 · 2018
Cited alongside, same era.
Reducing gender bias in abusive language detection
Ji Ho Park, Jamin Shin, and Pascale Fung. 2018 · 2018
Cited alongside, same era.
User-level race and ethnicity predictors from twitter text
Daniel Preoţiuc-Pietro and Lyle Ungar. 2018 · 2018
Cited alongside, same era.
Going beyond hate speech: The pragmatics of ethnic slur terms
Björn Technau. 2018 · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
Black and banned: Who is free speech for?
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, R’emi Louf, Morgan Funtowicz, and Jamie Brew. 2019 · 2019
Later among the works it cites.
Language (technology) is power: A critical survey of “bias” in NLP
Su Lin Blodgett, Solon Barocas, Hal Daumé, III, and Hanna Wallach. 2020 · 2020
Later among the works it cites.
Adversarial filters of dataset biases
Ronan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers, Matthew Peters, Ashish Sabharwal, and Yejin Choi. 2020 · 2020
Later among the works it cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Danyaal Yasin. 2018 · 2018
Cited alongside, same era.
Mitigating unwanted biases with adversarial learning
Brian Hu Zhang, Blake Lemoine, and Margaret Mitchell. 2018 · 2018
Cited alongside, same era.
Don’t take the easy way out: Ensemble based methods for avoiding known dataset biases
Christopher Clark, Mark Yatskar, and Luke Zettlemoyer. 2019 · 2019
Cited alongside, same era.
Racial bias in hate speech and abusive language detection datasets
Thomas Davidson, Debasmita Bhattacharya, and Ingmar Weber. 2019 · 2019
Cited alongside, same era.
Build it break it fix it for dialogue safety: Robustness from adversarial human attack
Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston. 2019 · 2019
Cited alongside, same era.
Unlearn dataset bias in natural language inference by fitting the residual
He He, Sheng Zha, and Haohan Wang. 2019 · 2019
Cited alongside, same era.
Behind the screen: Content moderation in the shadows of social media
Sarah T Roberts. 2019 · 2019
Cited alongside, same era.
Fighting hate speech, silencing drag queens? artificial intelligence in content moderation and risks to lgbtq voices online
Thiago Dias Oliva, Dennys Marcelo Antonialli, and Alessandra Gomes. 2020 · 2020
Later among the works it cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020 · 2020
Later among the works it cites.
Investigating African-American vernacular english in Transformer-Based text generation
Sophie Groenwold, Lily Ou, Aesha Parekh, Samhita Honnavalli, Sharon Levy, Diba Mirza, and William Yang Wang. 2020 · 2020
Later among the works it cites.
What civil rights groups want from facebook boycott: Stop hate speech and harassment of black users
Jessica Guynn. 2020 · 2020
Later among the works it cites.
Lessons from archives: strategies for collecting sociocultural data in machine learning
Eun Seo Jo and Timnit Gebru. 2020 · 2020
Later among the works it cites.
End-to-end bias mitigation by modelling biases in corpora
Rabeeh Karimi Mahabadi, Yonatan Belinkov, and James Henderson. 2020 · 2020
Later among the works it cites.
Intersectional bias in hate speech and abusive language datasets
Jae Yeon Kim, Carlos Ortiz, Sarah Nam, Sarah Santiago, and Vivek Datta. 2020 · 2020
Later among the works it cites.
Hate speech detection and racial bias mitigation in social media based on bert model
Marzieh Mozafari, Reza Farahbakhsh, and Noël Crespi. 2020 · 2020
Later among the works it cites.
Social bias frames: Reasoning about social and power implications of language
Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A Smith, and Yejin Choi. 2020 · 2020
Later among the works it cites.
Dataset cartography: Mapping and diagnosing datasets with training dynamics
Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, and Yejin Choi. 2020 · 2020
Later among the works it cites.
Mind the trade-off: Debiasing NLU models without degrading the in-distribution performance
Prasetya Ajie Utama, Nafise Sadat Moosavi, and Iryna Gurevych. 2020 · 2020
Later among the works it cites.
Demoting racial bias in hate speech detection
Mengzhou Xia, Anjalie Field, and Yulia Tsvetkov. 2020 · 2020
Later among the works it cites.