Fetching the paper…
Reading the bibliography…
Hate speech classifiers trained on imbalanced datasets struggle to determine if group identifiers like "gay" or "black" are used in offensive or prejudiced ways.
Incorporating priors with feature attribution on text classification
Frederick Liu and Besim Avci. 2019 · 1906
Earlier work this paper cites.
Interpretations are useful: penalizing explanations to align neural networks with prior knowledge
Laura Rieger, Chandan Singh, W James Murdoch, and Bin Yu. 2019 · 1909
Earlier work this paper cites.
Hate speech or “reasonable racism?” the other in stormfront
Priscilla Marie Meddaugh and Jack Kay. 2009 · 2009
Earlier work this paper cites.
The harm in hate speech
Jeremy Waldron. 2012 · 2012
Earlier work this paper cites.
Detecting hate speech on the world wide web
William Warner and Julia Hirschberg. 2012 · 2012
Earlier work this paper cites.
Countering online hate speech
Iginio Gagliardone, Danit Gal, Thiago Alves, and Gabriela Martinez. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Inside the hate-filled echo chamber of racism and conspiracy theories
Andrew Anthony. 2016 · 2016
Earlier work this paper cites.
Inside the twitter for racists: Gab the site where milo yiannopoulos goes to troll now
Thor Benson. 2016 · 2016
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. 2016 · 2016
Earlier work this paper cites.
Evidencing the harms of hate speech
Katharine Gelber and Luke McNamara. 2016 · 2016
Earlier work this paper cites.
Why should i trust you?: Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Analyzing the targets of hate in online social media
Leandro Silva, Mainack Mondal, Denzil Correa, Fabrício Benevenuto, and Ingmar Weber. 2016 · 2016
Earlier work this paper cites.
Hateful symbols or hateful people? predictive features for hate speech detection on twitter
Zeerak Waseem and Dirk Hovy. 2016 · 2016
Earlier work this paper cites.
Automated hate speech detection and the problem of offensive language
Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017 · 2017
Earlier work this paper cites.
Hate me, hate me not: Hate speech detection on facebook
Fabio Del Vigna12, Andrea Cimino23, Felice Dell’Orletta, Marinella Petrocchi, and Maurizio Tesconi. 2017 · 2017
Cited alongside, same era.
The impact of toxic language on the health of reddit communities
Shruthi Mohan, Apala Guha, Michael Harris, Fred Popowich, Ashley Schuster, and Chris Priebe. 2017 · 2017
Cited alongside, same era.
A measurement study of hate speech in social media
Mainack Mondal, Leandro Araújo Silva, and Fabrício Benevenuto. 2017 · 2017
Cited alongside, same era.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Understanding abuse: A typology of abusive language detection subtasks
Zeerak Waseem, Thomas Davidson, Dana Warmsley, and Ingmar Weber. 2017 · 2017
Semeval-2019 task 5: Multilingual detection of hate speech against immigrants and women in twitter
Valerio Basile, Cristina Bosco, Elisabetta Fersini, Debora Nozza, Viviana Patti, Francisco Manuel Rangel Pardo, Paolo Rosso, and Manuela Sanguinetti. 2019 · 2019
Later among the works it cites.
Counterfactual fairness in text classification through robustness
Sahaj Garg, Vincent Perot, Nicole Limtiaco, Ankur Taly, Ed H Chi, and Alex Beutel. 2019 · 2019
Later among the works it cites.
A survey of methods for explaining black box models
Riccardo Guidotti, Anna Monreale, Salvatore Ruggieri, Franco Turini, Fosca Giannotti, and Dino Pedreschi. 2019 · 2019
Later among the works it cites.
Bound in hatred: The role of group-based morality in acts of hate
Joseph Hoover, Mohammad Atari, Aida Mostafazadeh Davani, Brendan Kennedy, Gwenyth Portillo-Wightman, Leigh Yeh, Drew Kogon, and Morteza Dehghani. 2019 · 2019
Later among the works it cites.
Hate speech detection: Challenges and solutions
Sean MacAvaney, Hao-Ren Yao, Eugene Yang, Katina Russell, Nazli Goharian, and Ophir Frieder. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Ex machina: Personal attacks seen at scale
Ellery Wulczyn, Nithum Thain, and Lucas Dixon. 2017 · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Measuring and mitigating unintended bias in text classification
Lucas Dixon, John Li, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2018 · 2018
Cited alongside, same era.
Hate speech dataset from a white supremacy forum
Ona de Gibert, Naiara Perez, Aitor García Pablos, and Montse Cuadros. 2018 · 2018
Cited alongside, same era.
Inside the right-leaning echo chambers: Characterizing gab, an unmoderated social system
Lucas Lima, Julio CS Reis, Philipe Melo, Fabricio Murai, Leandro Araujo, Pantelis Vikatos, and Fabricio Benevenuto. 2018 · 2018
Cited alongside, same era.
Beyond word importance: Contextual decomposition to extract interactions from LSTMs
W. James Murdoch, Peter J. Liu, and Bin Yu. 2018 · 2018
Cited alongside, same era.
Overview of the hasoc track at fire 2019: Hate speech and offensive content identification in indo-european languages
Thomas Mandl, Sandip Modha, Prasenjit Majumder, Daksh Patel, Mohana Dave, Chintak Mandlia, and Aditya Patel. 2019 · 2019
Later among the works it cites.
Probing neural network comprehension of natural language arguments
Timothy Niven and Hung-Yu Kao. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
The risk of racial bias in hate speech detection
Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A Smith. 2019 · 2019
Later among the works it cites.
Hierarchical interpretations for neural network predictions
Chandan Singh, W. James Murdoch, and Bin Yu. 2019 · 2019
Later among the works it cites.
Detection of abusive language: the problem of biased datasets
Michael Wiegand, Josef Ruppenhofer, and Thomas Kleinbauer. 2019 · 2019
Later among the works it cites.
Towards hierarchical importance attribution: Explaining compositional semantics for neural sequence models
Xisen Jin, Zhongyu Wei, Junyi Du, Xiangyang Xue, and Xiang Ren. 2020 · 2020
Closest in time.
The gab hate corpus: A collection of 27k posts annotated for hate speech
Brendan Kennedy, Mohammad Atari, Aida M Davani, Leigh Yeh, Ali Omrani, Yehsong Kim, Kris Coombs Jr., Shreya Havaldar, Gwenyth Portillo-Wightman, Elaine Gonzalez, Joe Hoover, Aida Azatian, Gabriel Cardenas, Alyzeh Hussain, Austin Lara, Adam Omary, Christina Park, Xin Wang, Clarisa Wijaya, Yong Zhang, Beth Meyerowitz, and Morteza Dehghani. 2020 · 2020
Closest in time.
Hatred is in the eye of the annotator: Hate speech classifiers learn human-like social stereotypes (in press)
Aida Mostafazadeh Davani, Mohammad Atari, Brendan Kennedy, Shreya Havaldar, and Morteza Dehghani. 2020 · 2020
Closest in time.