Fetching the paper…
Reading the bibliography…
To support safety and inclusion in online communications, significant efforts in NLP research have been put towards addressing the problem of abusive content detection, commonly defined as a supervised classification task.
Tackling online abuse: A survey of automated abuse detection methods
Pushkar Mishra, Helen Yannakoudakis, and Ekaterina Shutova. 2019 · 1908
Earlier work this paper cites.
Judgment under uncertainty: Heuristics and biases
Amos Tversky and Daniel Kahneman. 1974 · 1974
Earlier work this paper cites.
Best-worst analysis
Jordan J. Louviere and George G. Woodworth. 1990 · 1990
Earlier work this paper cites.
Questions and Answers in Attitude Surveys: Experiments on Question Form, Wording, and Context
Stanley Presser and Howard Schuman. 1996 · 1996
Earlier work this paper cites.
Smokey: Automatic recognition of hostile messages
Ellen Spertus. 1997 · 1997
Earlier work this paper cites.
Response styles in marketing research: A cross-national investigation
Hans Baumgartner and Jan-Benedict E.M. Steenkamp. 2001 · 2001
Earlier work this paper cites.
Stereoset: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy. 2020 · 2004
Earlier work this paper cites.
Directions in abusive language training data: Garbage in, garbage out
Bertie Vidgen and Leon Derczynski. 2020 · 2004
Earlier work this paper cites.
Operationalizing the legal concept of ‘incitement to hatred’ as an NLP task
Frederike Zufall, Huangpan Zhang, Katharina Kloppenborg, and Torsten Zesch. 2020 · 2004
Earlier work this paper cites.
Multi-dimensional gender bias classification
Emily Dinan, Angela Fan, Ledell Wu, Jason Weston, Douwe Kiela, and Adina Williams. 2020 · 2005
Earlier work this paper cites.
Systematic attack surface reduction for deployed sentiment analysis models
Josh Kalin, David Noever, and Gerry Dozier. 2020 · 2006
Earlier work this paper cites.
Offensive language detection using multi-level classification
Amir H Razavi, Diana Inkpen, Sasha Uritsky, and Stan Matwin. 2010 · 2010
Earlier work this paper cites.
Is it all relative? Comparative judgments and the possible improvement of self-ratings and ratings of others
Richard D. Goffin and James M. Olson. 2011 · 2011
Earlier work this paper cites.
Improving cyberbullying detection with user context
Maral Dadvar, Dolf Trieschnigg, Roeland Ordelman, and Franciska de Jong. 2013 · 2013
Earlier work this paper cites.
Best-worst scaling: theory and methods
T. N. Flynn and A. A. J. Marley. 2014 · 2014
Earlier work this paper cites.
Hate speech detection with comment embeddings
Nemanja Djuric, Jing Zhou, Robin Morris, Mihajlo Grbovic, Vladan Radosavljevic, and Narayan Bhamidipati. 2015 · 2015
Earlier work this paper cites.
A lexicon-based approach for hate speech detection
Njagi Dennis Gitari, Zhang Zuping, Hanyurwimfura Damien, and Jun Long. 2015 · 2015
Earlier work this paper cites.
Analyzing labeled cyberbullying incidents on the Instagram social network
Homa Hosseinmardi, Sabrina Arredondo Mattson, Rahat Ibn Rafiq, Richard Han, Qin Lv, and Shivakant Mishra. 2015 · 2015
Earlier work this paper cites.
Best-Worst Scaling: Theory, Methods and Applications
Jordan J. Louviere, Terry N. Flynn, and A. A. J. Marley. 2015 · 2015
Earlier work this paper cites.
Histories of hating
Tamara Shepherd, Alison Harvey, Tim Jordan, Sam Srauy, and Kate Miltner. 2015 · 2015
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. 2016 · 2016
Earlier work this paper cites.
Capturing reliable fine-grained sentiment associations by crowdsourcing and best–worst scaling
Svetlana Kiritchenko and Saif M. Mohammad. 2016 · 2016
Earlier work this paper cites.
Abusive language detection in online user content
Chikashi Nobata, Joel Tetreault, Achint Thomas, Yashar Mehdad, and Yi Chang. 2016 · 2016
Earlier work this paper cites.
“Why should I trust you?” Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Commercial content moderation: Digital laborers’ dirty work
Sarah T. Roberts. 2016 · 2016
Earlier work this paper cites.
Are you a racist or am I seeing things? Annotator influence on hate speech detection on Twitter
Zeerak Waseem. 2016 · 2016
Earlier work this paper cites.
Hateful symbols or hateful people? Predictive features for hate speech detection on Twitter
Zeerak Waseem and Dirk Hovy. 2016 · 2016
Earlier work this paper cites.
Classification and its consequences for online harassment: Design insights from HeartMob
Lindsay Blackwell, Jill Dimond, Sarita Schoenebeck, and Cliff Lampe. 2017 · 2017
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017 · 2017
Earlier work this paper cites.
Mean birds: Detecting aggression and bullying on Twitter
Despoina Chatzakou, Nicolas Kourtellis, Jeremy Blackburn, Emiliano De Cristofaro, Gianluca Stringhini, and Athena Vakali. 2017 · 2017
Earlier work this paper cites.
Automated hate speech detection and the problem of offensive language
Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017 · 2017
Earlier work this paper cites.
Online Harassment 2017
Maeve Duggan. 2017 · 2017
Cited alongside, same era.
Legal framework, dataset and annotation schema for socially unacceptable online discourse practices in Slovene
Darja Fišer, Tomaž Erjavec, and Nikola Ljubešić. 2017 · 2017
Cited alongside, same era.
Detecting online hate speech using context aware models
Lei Gao and Ruihong Huang. 2017 · 2017
Cited alongside, same era.
A large labeled corpus for online harassment research
Jennifer Golbeck, Zahra Ashktorab, Rashad O Banjo, Alexandra Berlinger, Siddharth Bhagwan, Cody Buntain, Paul Cheakalos, Alicia A Geller, Rajesh Kumar Gnanasekaran, Raja Rajan Gunasekaran, et al. 2017 · 2017
Cited alongside, same era.
Deceiving Google’s Perspective API built for detecting toxic comments
Hossein Hosseini, Sreeram Kannan, Baosen Zhang, and Radha Poovendran. 2017 · 2017
Cited alongside, same era.
Pay “attention” to your context when classifying abusive language
Tuhin Chakrabarty, Kilol Gupta, and Smaranda Muresan. 2019 · 2019
Later among the works it cites.
Racial bias in hate speech and abusive language detection datasets
Thomas Davidson, Debasmita Bhattacharya, and Ingmar Weber. 2019 · 2019
Later among the works it cites.
A unified deep learning architecture for abuse detection
Antigoni Maria Founta, Despoina Chatzakou, Nicolas Kourtellis, Jeremy Blackburn, Athena Vakali, and Ilias Leontiadis. 2019 · 2019
Later among the works it cites.
Counterfactual fairness in text classification through robustness
Sahaj Garg, Vincent Perot, Nicole Limtiaco, Ankur Taly, Ed H Chi, and Alex Beutel. 2019 · 2019
Later among the works it cites.
Attention is not Explanation
Sarthak Jain and Byron C. Wallace. 2019 · 2019
Later among the works it cites.
A just and comprehensive strategy for using NLP to address online abuse
David Jurgens, Libby Hemphill, and Eshwar Chandrasekharan. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jigsaw, Perspective API. 2017 · 2017
Cited alongside, same era.
Best-worst scaling more reliable than rating scales: A case study on sentiment intensity annotation
Svetlana Kiritchenko and Saif M. Mohammad. 2017 · 2017
Cited alongside, same era.
Extreme speech online: An anthropological critique of hate speech debates
Matti Pohjonen and Sahana Udupa. 2017 · 2017
Cited alongside, same era.
Hate speech annotation: Analysis of an Italian Twitter corpus
Fabio Poletto, Marco Stranisci, Manuela Sanguinetti, Viviana Patti, and Cristina Bosco. 2017 · 2017
Cited alongside, same era.
A survey on hate speech detection using natural language processing
Anna Schmidt and Michael Wiegand. 2017 · 2017
Cited alongside, same era.
Understanding abuse: A typology of abusive language detection subtasks
Zeerak Waseem, Thomas Davidson, Dana Warmsley, and Ingmar Weber. 2017 · 2017
Cited alongside, same era.
Ex machina: Personal attacks seen at scale
Ellery Wulczyn, Nithum Thain, and Lucas Dixon. 2017 · 2017
Cited alongside, same era.
Later among the works it cites.
Measuring bias in contextualized word representations
Keita Kurita, Nidhi Vyas, Ayush Pareek, Alan W Black, and Yulia Tsvetkov. 2019 · 2019
Later among the works it cites.
Black is to criminal as caucasian is to police: Detecting and removing multiclass bias in word embeddings
Thomas Manzini, Lim Yao Chong, Alan W Black, and Yulia Tsvetkov. 2019 · 2019
Later among the works it cites.
On measuring social biases in sentence encoders
Chandler May, Alex Wang, Shikha Bordia, Samuel Bowman, and Rachel Rudinger. 2019 · 2019
Later among the works it cites.
Model cards for model reporting
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019 · 2019
Later among the works it cites.
What is abusive language?
Marco Niemann, Dennis M Riehle, Jens Brunk, and Jörg Becker. 2019 · 2019
Later among the works it cites.
The risk of racial bias in hate speech detection
Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A Smith. 2019 · 2019
Later among the works it cites.
Detecting aggression and toxicity using a multi dimension capsule network
Saurabh Srivastava and Prerna Khurana. 2019 · 2019
Later among the works it cites.
Challenges and frontiers in abusive content detection
Bertie Vidgen, Alex Harris, Dong Nguyen, Rebekah Tromble, Scott Hale, and Helen Margetts. 2019 · 2019
Later among the works it cites.
Detection of abusive language: the problem of biased datasets
Michael Wiegand, Josef Ruppenhofer, and Thomas Kleinbauer. 2019 · 2019
Later among the works it cites.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter. 2019 · 2019
Later among the works it cites.
Gendered morality and backlash effects in online discussions: An experimental study on how users respond to hate speech comments against women and sexual minorities
Claudia Wilhelm and Sven Joeckel. 2019 · 2019
Later among the works it cites.
Semeval-2019 task 6: Identifying and categorizing offensive language in social media (offenseval)
Marcos Zampieri, Shervin Malmasi, Preslav Nakov, Sara Rosenthal, Noura Farra, and Ritesh Kumar. 2019b · 2019
Later among the works it cites.
Annotating for hate speech: The MaNeCo corpus and some input from critical discourse analysis
Stavros Assimakopoulos, Rebecca Vella Muskat, Lonneke van der Plas, and Albert Gatt. 2020 · 2020
Closest in time.
Language (technology) is power: A critical survey of “bias” in NLP
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020 · 2020
Closest in time.
Toxic, hateful, offensive or abusive? what are we really classifying? an empirical analysis of hate speech datasets
Paula Fortuna, Juan Soler, and Leo Wanner. 2020 · 2020
Closest in time.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Sam Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020 · 2020
Closest in time.
Cyberbullying fact sheet: Identification, Prevention, and Response
S. Hinduja and J. W. Patchin. 2020 · 2020
Closest in time.
Weight poisoning attacks on pre-trained models
Keita Kurita, Paul Michel, and Graham Neubig. 2020 · 2020
Closest in time.
Explainable AI approach towards toxic comment classification
Aditya Mahajan, Divyank Shah, and Gibraan Jafar. 2020 · 2020
Closest in time.
Crows-pairs: A challenge dataset for measuring social biases in masked language models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020 · 2020
Closest in time.
On cross-dataset generalization in automatic detection of online abuse
Isar Nejadgholi and Svetlana Kiritchenko. 2020 · 2020
Closest in time.
Approaches to automated detection of cyberbullying: A survey
S. Salawu, Y. He, and J. Lumsden. 2020 · 2020
Closest in time.
Social bias frames: Reasoning about social and power implications of language
Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A Smith, and Yejin Choi. 2020 · 2020
Closest in time.
A multi-platform dataset for detecting cyberbullying in social media
David Van Bruwaene, Qianjia Huang, and Diana Inkpen. 2020 · 2020
Closest in time.