Deep learning models for multilingual hate speech detection
Aluru, S. S., Mathew, B., Saha, P., and Mukherjee, A. (2020) · 2020
Closest in time.
A novel methodology for developing automatic harassment classifiers for Twitter
Arora, I., Guo, J., Levitan, S. I., McGregor, S., and Hirschberg, J. (2020) · 2020
Closest in time.
Annotating for hate speech: The MaNeCo corpus and some input from critical discourse analysis
Assimakopoulos, S., Vella Muskat, R., van der Plas, L., and Gatt, A. (2020) · 2020
Closest in time.
Machine learning techniques for hate speech classification of Twitter data: State-of-the-art, future challenges and research directions
Ayo, F. E., Folorunso, O., Ibharalu, F. T., and Osinuga, I. A. (2020) · 2020
Closest in time.
A unified taxonomy of harmful content
Banko, M., MacKeen, B., and Ray, L. (2020) · 2020
Closest in time.
Unmasking contextual stereotypes: Measuring and mitigating BERT’s gender bias
Bartl, M., Nissim, M., and Gatt, A. (2020) · 2020
Closest in time.
Language (technology) is power: A critical survey of “bias” in NLP
Blodgett, S. L., Barocas, S., Daumé III, H., and Wallach, H. (2020) · 2020
Closest in time.
WAC: A corpus of Wikipedia conversations for online abuse detection
Cecillon, N., Labatut, V., Dufour, R., and Linares, G. (2020) · 2020
Closest in time.
Multi-dimensional gender bias classification
Dinan, E., Fan, A., Wu, L., Weston, J., Kiela, D., and Williams, A. (2020) · 2020
Closest in time.
Towards transparency by design for artificial intelligence
Felzmann, H., Fosch-Villaronga, E., Lutz, C., and Tamò-Larrieux, A. (2020) · 2020
Closest in time.
Unsupervised discovery of implicit gender bias
Field, A., and Tsvetkov, Y. (2020) · 2020
Closest in time.
Principled artificial intelligence: Mapping consensus in ethical and rights-based approaches to principles for AI
Fjeld, J., Achten, N., Hilligoss, H., Nagy, A., and Srikumar, M. (2020) · 2020
Closest in time.
Toxic, hateful, offensive or abusive? What are we really classifying? An empirical analysis of hate speech datasets
Fortuna, P., Soler, J., and Wanner, L. (2020) · 2020
Closest in time.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A. (2020) · 2020
Closest in time.
Cyberbullying detection with fairness constraints
Gencoglu, O. (2020) · 2020
Closest in time.
Algorithmic content moderation: Technical and political challenges in the automation of platform governance
Gorwa, R., Binns, R., and Katzenbach, C. (2020) · 2020
Closest in time.
Why attention is not explanation: Surgical intervention and causal reasoning about neural models
Grimsley, C., Mayfield, E., and Bursten, J. R. (2020) · 2020
Closest in time.
Transformers and data augmentation for aggressiveness detection in Mexican Spanish
Guzman-Silverio, M., Balderas-Paredes, A., and López-Monroy, A. (2020) · 2020
Closest in time.
Gaming algorithmic hate-speech detection: Stakes, parties, and moves
Haapoja, J., Laaksonen, S.-M., and Lampinen, A. (2020) · 2020
Closest in time.
The political power of platforms: How current attempts to regulate misinformation amplify opinion power
Helberger, N. (2020) · 2020
Closest in time.
Cyberbullying fact sheet: Identification, Prevention, and Response
Hinduja, S., and Patchin, J. W. (2020) · 2020
Closest in time.
Multilingual Twitter corpus and baselines for evaluating demographic bias in hate speech recognition
Huang, X., Xing, L., Dernoncourt, F., and Paul, M. (2020) · 2020
Closest in time.
Lawmakers drill down on how facebook and twitter moderate content
Isaac, M., and Browning, K. (2020) · 2020
Closest in time.
What makes discrimination morally wrong? A harm-based view reconsidered
Ishida, S. (2020) · 2020
Closest in time.
The state and fate of linguistic diversity and inclusion in the NLP world
Joshi, P., Santy, S., Budhiraja, A., Bali, K., and Choudhury, M. (2020) · 2020
Closest in time.
Systematic attack surface reduction for deployed sentiment analysis models
Kalin, J., Noever, D., and Dozier, G. (2020) · 2020
Closest in time.
Toward situated interventions for algorithmic equity: lessons from the field
Katell, M., Young, M., Dailey, D., Herman, B., Guetler, V., Tam, A., Bintz, C., Raz, D., and Krafft, P. M. (2020) · 2020
Closest in time.
The hateful memes challenge: Detecting hate speech in multimodal memes
Kiela, D., Firooz, H., Mohan, A., Goswami, V., Singh, A., Ringshia, P., and Testuggine, D. (2020) · 2020
Closest in time.
Report: Facebook makes 300,000 content moderation mistakes every day
Koetsier, J. (2020) · 2020
Closest in time.
Weight poisoning attacks on pre-trained models
Kurita, K., Michel, P., and Neubig, G. (2020) · 2020
Closest in time.
ALBERT: A lite BERT for self-supervised learning of language representations
Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R. (2020) · 2020
Closest in time.
Explainable AI approach towards toxic comment classification
Mahajan, A., Shah, D., and Jafar, G. (2020) · 2020
Closest in time.
Predicting the outbreak of conflict in online discussions using emotion-based features
Marcinowski, M., and Ławrynowicz, A. (2020) · 2020
Closest in time.
Swe2: Subword enriched and significant word emphasized framework for hate speech detection
Mou, G., Ye, P., and Lee, K. (2020) · 2020
Closest in time.
Crows-pairs: A challenge dataset for measuring social biases in masked language models
Nangia, N., Vania, C., Bhalerao, R., and Bowman, S. R. (2020) · 2020
Closest in time.
A survey of pre-processing techniques to improve short-text quality: a case study on hate speech detection on Twitter
Naseem, U., Razzak, I., and Eklund, P. W. (2020) · 2020
Closest in time.
On cross-dataset generalization in automatic detection of online abuse
Nejadgholi, I., and Kiritchenko, S. (2020) · 2020
Closest in time.
Comparative evaluation of label agnostic selection bias in multilingual hate speech datasets
Ousidhoum, N., Song, Y., and Yeung, D.-Y. (2020) · 2020
Closest in time.
Toxicity detection: Does context really matter?
Pavlopoulos, J., Sorensen, J., Dixon, L., Thain, N., and Androutsopoulos, I. (2020) · 2020
Closest in time.
Monitoring users behavior: Anti-immigration speech detection on Twitter
Pitropakis, N., Kokot, K., Gkatzia, D., Ludwiniak, R., Mylonas, A., and Kandias, M. (2020) · 2020
Closest in time.
Resources and benchmark corpora for hate speech detection: a systematic review
Poletto, F., Basile, V., Sanguinetti, M., Bosco, C., and Patti, V. (2020) · 2020
Closest in time.
Online abuse and human rights: WOAH satellite session at RightsCon 2020
Prabhakaran, V., Waseem, Z., Akiwowo, S., and Vidgen, B. (2020) · 2020
Closest in time.
Solidarity and community engagement in global health research
Pratt, B., Cheah, P. Y., and Marsh, V. (2020) · 2020
Closest in time.
Six attributes of unhealthy conversations
Price, I., Gifford-Moore, J., Flemming, J., Musker, S., Roichman, M., Sylvain, G., Thain, N., Dixon, L., and Sorensen, J. (2020) · 2020
Closest in time.
Mitigating unfair bias in ML models with the mindiff framework
Prost, F. (2020) · 2020
Closest in time.
Closing the AI accountability gap: defining an end-to-end framework for internal algorithmic auditing
Raji, I. D., Smart, A., White, R. N., Mitchell, M., Gebru, T., Hutchinson, B., Smith-Loud, J., Theron, D., and Barnes, P. (2020) · 2020
Closest in time.
Investigating sampling bias in abusive language detection
Razo, D., and Kübler, S. (2020) · 2020
Closest in time.
Bagging BERT models for robust aggression identification
Risch, J., and Krestel, R. (2020) · 2020
Closest in time.
Aggression and misogyny detection using BERT: A multi-task approach
Safi Samghabadi, N., Patwa, P., PYKL, S., Mukherjee, P., Das, A., and Solorio, T. (2020) · 2020
Closest in time.
Approaches to automated detection of cyberbullying: A survey
Salawu, S., He, Y., and Lumsden, J. (2020) · 2020
Closest in time.
Developing an online hate classifier for multiple social media platforms
Salminen, J., Hopf, M., Chowdhury, S. A., Jung, S.-g., Almerekhi, H., and Jansen, B. J. (2020) · 2020
Closest in time.
Social bias frames: Reasoning about social and power implications of language
Sap, M., Gabriel, S., Qin, L., Jurafsky, D., Smith, N. A., and Choi, Y. (2020) · 2020
Closest in time.
What happened when humans stopped managing social media content
Scott, M., and Kayali, L. (2020) · 2020
Closest in time.
Predictive biases in natural language processing models: A conceptual framework and overview
Shah, D. S., Schwartz, H. A., and Hovy, D. (2020) · 2020
Closest in time.
Offensive language and hate speech detection for Danish
Sigurbergsson, G. I., and Derczynski, L. (2020) · 2020
Closest in time.
Generating counter narratives against online hate speech: Data and strategies
Tekiroğlu, S. S., Chung, Y.-L., and Guerini, M. (2020) · 2020
Closest in time.
Detecting ‘dirt and ‘toxicity: Rethinking content moderation as pollution behaviour
Thylstrup, N., and Waseem, Z. (2020) · 2020
Closest in time.
Towards a friendly online community: An unsupervised style transfer framework for profanity redaction
Tran, M., Zhang, Y., and Soleymani, M. (2020) · 2020
Closest in time.
Dispute resolution and content moderation: Fair, accountable, independent, transparent, and effective
Tworek, H., and Tenove, C. (2020) · 2020
Closest in time.
Quarantining online hate speech: technical and ethical perspectives
Ullmann, S., and Tomalin, M. (2020) · 2020
Closest in time.
A multi-platform dataset for detecting cyberbullying in social media
Van Bruwaene, D., Huang, Q., and Inkpen, D. (2020) · 2020
Closest in time.
On mismatched detection and safe, trustworthy machine learning
Varshney, K. R. (2020) · 2020
Closest in time.
A human-centered agenda for intelligible machine learning
Vaughan, J. W., and Wallach, H. (2020) · 2020
Closest in time.
Directions in abusive language training data, a systematic review: Garbage in, garbage out
Vidgen, B., and Derczynski, L. (2020) · 2020
Closest in time.
Investigating annotator bias with a graph-based approach
Wich, M., Al Kuwatly, H., and Groh, G. (2020) · 2020
Closest in time.
UHH-LT at SemEval-2020 task 12: Fine-tuning of pre-trained transformer networks for offensive language detection
Wiedemann, G., Yimam, S. M., and Biemann, C. (2020) · 2020
Closest in time.
Towards hate speech detection at large via deep generative modeling
Wullach, T., Adler, A., and Minkov, E. M. (2020) · 2020
Closest in time.
SemEval-2020 task 12: Multilingual offensive language identification in social media (OffensEval 2020)
Zampieri, M., Nakov, P., Rosenthal, S., Atanasova, P., Karadzhov, G., Mubarak, H., Derczynski, L., Pitenis, Z., and Çöltekin, Ç. (2020) · 2020
Closest in time.
Understanding and countering stereotypes: A computational approach to the stereotype content model
Fraser, K. C., Nejadgholi, I., and Kiritchenko, S. (2021) · 2021
Closest in time.
On transferability of bias mitigation effects in language model fine-tuning
Jin, X., Barbieri, F., Kennedy, B., Davani, A. M., Neves, L., and Ren, X. (2021) · 2021
Closest in time.
HateCheck: Functional tests for hate speech detection models
Röttger, P., Vidgen, B., Nguyen, D., Waseem, Z., Margetts, H., and Pierrehumbert, J. (2021) · 2021
Closest in time.
Introducing CAD: the contextual abuse dataset
Vidgen, B., Nguyen, D., Margetts, H., Rossini, P., and Tromble, R. (2021) · 2021
Closest in time.
Implicitly abusive language – what does it actually look like and why are we not getting there?
Wiegand, M., Ruppenhofer, J., and Eder, E. (2021) · 2021
Closest in time.
Challenges in automated debiasing for toxic language detection
Zhou, X., Sap, M., Swayamdipta, S., Choi, Y., and Smith, N. (2021) · 2021
Closest in time.