Fetching the paper…
Reading the bibliography…
Detecting online hate is a difficult task that even state-of-the-art models struggle with.
Tackling online abuse: A survey of automated abuse detection methods
Pushkar Mishra, Helen Yannakoudakis, and Ekaterina Shutova. 2020 · 1908
Earlier work this paper cites.
The measurement of observer agreement for categorical data
J. Richard Landis and Gary G. Koch. 1977 · 1977
Earlier work this paper cites.
Grounded theory research: Procedures, canons, and evaluative criteria
Juliet M Corbin and Anselm Strauss. 1990 · 1990
Earlier work this paper cites.
Black-box testing: techniques for functional testing of software and systems
Boris Beizer. 1995 · 1995
Earlier work this paper cites.
Unsex me here: Revisiting sexism detection using psychological scales and adversarial samples
Mattia Samory, Indira Sen, Julian Kohne, Fabian Floeck, and Claudia Wagner. 2020 · 2004
Earlier work this paper cites.
Learning from imbalanced data
Haibo He and Edwardo A Garcia. 2009 · 2009
Earlier work this paper cites.
Computing inter-rater reliability for observational data: An overview and tutorial
Kevin A. Hallgren. 2012 · 2012
Earlier work this paper cites.
The Winograd schema challenge
Hector Levesque, Ernest Davis, and Leora Morgenstern. 2012 · 2012
Earlier work this paper cites.
Learning from the worst: Dynamically generated datasets to improve online hate detection
Bertie Vidgen, Tristan Thrush, Zeerak Waseem, and Douwe Kiela. 2020b · 2012
Earlier work this paper cites.
Detecting hate speech on the World Wide Web
William Warner and Julia Hirschberg. 2012 · 2012
Earlier work this paper cites.
Cyber hate speech on Twitter: An application of machine classification and statistical modeling for policy and decision making
Pete Burnap and Matthew L Williams. 2015 · 2015
Earlier work this paper cites.
Abusive language detection in online user content
Chikashi Nobata, Joel Tetreault, Achint Thomas, Yashar Mehdad, and Yi Chang. 2016 · 2016
Earlier work this paper cites.
Are you a racist or am I seeing things? Annotator influence on hate speech detection on Twitter
Zeerak Waseem. 2016 · 2016
Earlier work this paper cites.
Hateful symbols or hateful people? Predictive features for hate speech detection on Twitter
Zeerak Waseem and Dirk Hovy. 2016 · 2016
Earlier work this paper cites.
Automated hate speech detection and the problem of offensive language
Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017 · 2017
Earlier work this paper cites.
A large labeled corpus for online harassment research
Jennifer Golbeck, Zahra Ashktorab, Rashad O Banjo, Alexandra Berlinger, Siddharth Bhagwan, Cody Buntain, Paul Cheakalos, Alicia A Geller, Rajesh Kumar Gnanasekaran, Raja Rajan Gunasekaran, et al. 2017 · 2017
Earlier work this paper cites.
Deceiving Google’s Perspective API built for detecting toxic comments
Hossein Hosseini, Sreeram Kannan, Baosen Zhang, and Radha Poovendran. 2017 · 2017
Earlier work this paper cites.
A challenge set approach to evaluating machine translation
Pierre Isabelle, Colin Cherry, and George Foster. 2017 · 2017
Earlier work this paper cites.
Systems and Software Engineering – Vocabulary
ISO/IEC/IEEE 24765:2017(E). 2017 · 2017
Earlier work this paper cites.
A survey on hate speech detection using natural language processing
Anna Schmidt and Michael Wiegand. 2017 · 2017
Earlier work this paper cites.
How grammatical is character-level neural machine translation? Assessing MT quality with contrastive translation pairs
Rico Sennrich. 2017 · 2017
Earlier work this paper cites.
Understanding abuse: A typology of abusive language detection subtasks
Zeerak Waseem, Thomas Davidson, Dana Warmsley, and Ingmar Weber. 2017 · 2017
Earlier work this paper cites.
Ex machina: Personal attacks seen at scale
Ellery Wulczyn, Nithum Thain, and Lucas Dixon. 2017 · 2017
Earlier work this paper cites.
Challenges for toxic comment classification: An in-depth error analysis
Betty van Aken, Julian Risch, Ralf Krestel, and Alexander Löser. 2018 · 2018
Earlier work this paper cites.
Synthetic and natural noise both break neural machine translation
Yonatan Belinkov and Yonatan Bisk. 2018 · 2018
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M. Bender and Batya Friedman. 2018 · 2018
Cited alongside, same era.
Measuring and mitigating unintended bias in text classification
Lucas Dixon, John Li, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2018 · 2018
Cited alongside, same era.
A survey on automatic detection of hate speech in text
Paula Fortuna and Sérgio Nunes. 2018 · 2018
Cited alongside, same era.
Large scale crowdsourcing and characterization of Twitter abusive behavior
Antigoni Maria Founta, Constantinos Djouvas, Despoina Chatzakou, Ilias Leontiadis, Jeremy Blackburn, Gianluca Stringhini, Athena Vakali, Michael Sirivianos, and Nicolas Kourtellis. 2018 · 2018
Cited alongside, same era.
Breaking NLI systems with sentences that require simple lexical inferences
Max Glockner, Vered Shwartz, and Yoav Goldberg. 2018 · 2018
Cited alongside, same era.
Detection of abusive language: The problem of biased datasets
Michael Wiegand, Josef Ruppenhofer, and Thomas Kleinbauer. 2019 · 2019
Later among the works it cites.
Errudite: Scalable, reproducible, and testable error analysis
Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel Weld. 2019 · 2019
Later among the works it cites.
Predicting the type and target of offensive posts in social media
Marcos Zampieri, Shervin Malmasi, Preslav Nakov, Sara Rosenthal, Noura Farra, and Ritesh Kumar. 2019 · 2019
Later among the works it cites.
Hate speech detection: A solved problem? The challenging case of long tail on Twitter
Ziqi Zhang and Lei Luo. 2019 · 2019
Later among the works it cites.
A Unified Taxonomy of Harmful Content
Michele Banko, Brendon MacKeen, and Laurie Ray. 2020 · 2020
Closest in time.
DeepHate: Hate speech detection via multi-faceted text representations
Rui Cao, Roy Ka-Wei Lee, and Tuan-Anh Hoang. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tommi Gröndahl, Luca Pajola, Mika Juuti, Mauro Conti, and N Asokan. 2018 · 2018
Cited alongside, same era.
Challenges in discriminating profanity from hate speech
Shervin Malmasi and Marcos Zampieri. 2018 · 2018
Cited alongside, same era.
Targeted syntactic evaluation of language models
Rebecca Marvin and Tal Linzen. 2018 · 2018
Cited alongside, same era.
Stress test evaluation for natural language inference
Aakanksha Naik, Abhilasha Ravichander, Norman Sadeh, Carolyn Rose, and Graham Neubig. 2018 · 2018
Cited alongside, same era.
Reducing gender bias in abusive language detection
Ji Ho Park, Jamin Shin, and Pascale Fung. 2018 · 2018
Cited alongside, same era.
Leveraging intra-user and inter-user representation learning for automated hate speech detection
Jing Qian, Mai ElSherief, Elizabeth Belding, and William Yang Wang. 2018 · 2018
Cited alongside, same era.
Anatomy of online hate: Developing a taxonomy and machine learning models for identifying and classifying hate in online news media
Joni Salminen, Hind Almerekhi, Milica Milenković, Soon-gyo Jung, Jisun An, Haewoon Kwak, and Bernard Jansen. 2018 · 2018
Cited alongside, same era.
Closest in time.
What BERT is not: Lessons from a new suite of psycholinguistic diagnostics for language models
Allyson Ettinger. 2020 · 2020
Closest in time.
Evaluating models’ local decision boundaries via contrast sets
Matt Gardner, Yoav Artzi, Victoria Basmov, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, Nitish Gupta, Hannaneh Hajishirzi, Gabriel Ilharco, Daniel Khashabi, Kevin Lin, Jiangming Liu, Nelson F. Liu, Phoebe Mulcaire, Qiang Ning, Sameer Singh, Noah A. Smith, Sanjay Subramanian, Reut Tsarfaty, Eric Wallace, Ally Zhang, and Ben Zhou. 2020 · 2020
Closest in time.
Learning the difference that makes a difference with counterfactually-augmented data
Divyansh Kaushik, Eduard Hovy, and Zachary C. Lipton. 2020 · 2020
Closest in time.
Contextualizing hate speech classifiers with post-hoc explanation
Brendan Kennedy, Xisen Jin, Aida Mostafazadeh Davani, Morteza Dehghani, and Xiang Ren. 2020 · 2020
Closest in time.
Towards a comprehensive taxonomy and large-scale annotated corpus for online slur usage
Jana Kurrek, Haji Mohammad Saleem, and Derek Ruths. 2020 · 2020
Closest in time.
A framework for the computational linguistic analysis of dehumanization
Julia Mendelsohn, Yulia Tsvetkov, and Dan Jurafsky. 2020 · 2020
Closest in time.
On cross-dataset generalization in automatic detection of online abuse
Isar Nejadgholi and Svetlana Kiritchenko. 2020 · 2020
Closest in time.
Adversarial NLI: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020 · 2020
Closest in time.
COLD: Annotation scheme and evaluation data set for complex offensive language in English
Alexis Palmer, Christine Carr, Melissa Robinson, and Jordan Sanders. 2020 · 2020
Closest in time.
Resources and benchmark corpora for hate speech detection: a systematic review
Fabio Poletto, Valerio Basile, Manuela Sanguinetti, Cristina Bosco, and Viviana Patti. 2020 · 2020
Closest in time.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Closest in time.
Social bias frames: Reasoning about social and power implications of language
Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. 2020 · 2020
Closest in time.
Predictive biases in natural language processing models: A conceptual framework and overview
Deven Santosh Shah, H. Andrew Schwartz, and Dirk Hovy. 2020 · 2020
Closest in time.
HABERTOR: An efficient and effective deep hatespeech detector
Thanh Tran, Yifan Hu, Changwei Hu, Kevin Yen, Fei Tan, Kyumin Lee, and Se Rim Park. 2020 · 2020
Closest in time.
Directions in abusive language training data, a systematic review: Garbage in, garbage out
Bertie Vidgen and Leon Derczynski. 2020 · 2020
Closest in time.
BLiMP: The benchmark of linguistic minimal pairs for English
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R. Bowman. 2020 · 2020
Closest in time.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Closest in time.
Dynabench: Rethinking benchmarking in nlp
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, et al. 2021 · 2021
Closest in time.