Fetching the paper…
Reading the bibliography…
Modern toxic speech detectors are incompetent in recognizing disguised offensive language, such as adversarial attacks that deliberately avoid known toxic lexicons, or manifestations of implicit bias.
HuggingFace’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, R’emi Louf, Morgan Funtowicz, and Jamie Brew. 2019 · 1910
Earlier work this paper cites.
A sequential algorithm for training text classifiers
David D. Lewis and William A. Gale. 1994 · 1994
Earlier work this paper cites.
Estimating training data influence by tracking gradient descent
Garima Pruthi, Frederick Liu, Mukund Sundararajan, and Satyen Kale. 2020 · 2002
Earlier work this paper cites.
Racial microaggressions in everyday life: implications for clinical practice
Derald Wing Sue, Christina M. Capodilupo, Gina C. Torino, Jennifer M Bucceri, Aisha M. B. Holder, Kevin L Nadal, and Marta Esquilin. 2007 · 2007
Earlier work this paper cites.
Microaggressions in everyday life: Race, gender, and sexual orientation
Derald Wing Sue. 2010 · 2010
Earlier work this paper cites.
The impact of racial microaggressions on mental health: Counseling implications for clients of color
Kevin L Nadal, Katie E Griffin, Yinglee Wong, Sahran Hamit, and Morgan Rasmus. 2014 · 2014
Earlier work this paper cites.
Hateful symbols or hateful people? predictive features for hate speech detection on twitter
Zeerak Waseem and Dirk Hovy. 2016 · 2016
Earlier work this paper cites.
Second-order stochastic optimization for machine learning in linear time
Naman Agarwal, Brian Bullins, and Elad Hazan. 2017 · 2017
Earlier work this paper cites.
Automated hate speech detection and the problem of offensive language
Thomas Davidson, Dana Warmsley, Michael W. Macy, and Ingmar Weber. 2017 · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang. 2017 · 2017
Cited alongside, same era.
A survey on hate speech detection using natural language processing
Anna Schmidt and Michael Wiegand. 2017 · 2017
Cited alongside, same era.
Surfacing contextual hate speech words within social media
Jherez Taylor, Melvyn Peignon, and Yi-Shin Chen. 2017 · 2017
Cited alongside, same era.
A survey on automatic detection of hate speech in text
Paula Fortuna and Sérgio Nunes. 2018 · 2018
Cited alongside, same era.
Large scale crowdsourcing and characterization of twitter abusive behavior
Antigoni-Maria Founta, Constantinos Djouvas, Despoina Chatzakou, Ilias Leontiadis, Jeremy Blackburn, Gianluca Stringhini, Athena Vakali, Michael Sirivianos, and Nicolas Kourtellis. 2018 · 2018
Cited alongside, same era.
A just and comprehensive strategy for using NLP to address online abuse
David Jurgens, Libby Hemphill, and Eshwar Chandrasekharan. 2019 · 2019
Later among the works it cites.
Interpreting black box predictions using fisher kernels
Rajiv Khanna, B. Kim, Joydeep Ghosh, and O. Koyejo. 2019 · 2019
Later among the works it cites.
Hate speech detection: Challenges and solutions
Sean MacAvaney, Hao-Ren Yao, Eugene Yang, Katina Russell, Nazli Goharian, and Ophir Frieder. 2019 · 2019
Later among the works it cites.
Talkdown: A corpus for condescension detection in context
Zijian Wang and Christopher Potts. 2019 · 2019
Later among the works it cites.
Learning a stopping criterion for active learning for word sense disambiguation and text classification
Jingbo Zhu, Huizhen Wang, and Eduard H. Hovy. 2008 · 2019
Later among the works it cites.
Relatif: Identifying explanatory training examples via relative influence
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Edwin Jain, Stephan Brown, Jeffery Chen, Erin Neaton, Mohammad Baidas, Ziqian Dong, Huanying Gu, and Nabi Sertac Artan. 2018 · 2018
Cited alongside, same era.
Representer point selection for explaining deep neural networks
Chih-Kuan Yeh, Joon Sik Kim, Ian En-Hsu Yen, and Pradeep Ravikumar. 2018 · 2018
Cited alongside, same era.
Finding microaggressions in the wild: A case for locating elusive phenomena in social media posts
Luke Breitfeller, Emily Ahn, David Jurgens, and Yulia Tsvetkov. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Elnaz Barshan, Marc-Etienne Brunet, and G. Dziugaite. 2020 · 2020
Closest in time.
Unsupervised discovery of implicit gender bias
Anjalie Field and Yulia Tsvetkov. 2020 · 2020
Closest in time.
Explaining black box predictions and unveiling data artifacts through influence functions
Xiaochuang Han, Byron C. Wallace, and Yulia Tsvetkov. 2020 · 2020
Closest in time.
Social bias frames: Reasoning about social and power implications of language
Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. 2020 · 2020
Closest in time.