Fetching the paper…
Reading the bibliography…
Automatic hate speech detection using deep neural models is hampered by the scarcity of labeled datasets, leading to poor generalization.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. (2020) · 1901
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019) · 1907
Earlier work this paper cites.
Hatebert: Retraining BERT for abusive language detection in english
Caselli, T., Basile, V., Mitrovic, J., and Granitzer, M. (2020) · 2010
Earlier work this paper cites.
Hateful symbols or hateful people? predictive features for hate speech detection on twitter
Waseem, Z. and Hovy, D. (2016) · 2016
Earlier work this paper cites.
Automated hate speech detection and the problem of offensive language
Davidson, T., Warmsley, D., Macy, M., and Weber, I. (2017) · 2017
Earlier work this paper cites.
Detecting online hate speech using context aware models
Gao, L. and Huang, R. (2017) · 2017
Earlier work this paper cites.
Hate speech dataset from a white supremacy forum
de Gibert, O., Perez, N., García-Pablos, A., and Cuadros, M. (2018) · 2018
Earlier work this paper cites.
A survey on automatic detection of hate speech in text
Fortuna, P. and Nunes, S. (2018) · 2018
Earlier work this paper cites.
Large scale crowdsourcing and characterization of twitter abusive behavior
Founta, A. M., Djouvas, C., Chatzakou, D., Leontiadis, I., Blackburn, J., Stringhini, G., Vakali, A., Sirivianos, M., and Kourtellis, N. (2018) · 2018
Earlier work this paper cites.
Semeval-2019 task 5: Multilingual detection of hate speech against immigrants and women in twitter
Basile, V., Bosco, C., Fersini, E., Nozza, D., Patti, V., Pardo, F. M. R., Rosso, P., and Sanguinetti, M. (2019) · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. (2019) · 2019
Earlier work this paper cites.
Detection of abusive language: the problem of biased datasets
Wiegand, M., Ruppenhofer, J., and Kleinbauer, T. (2019) · 2019
Earlier work this paper cites.
Do not have enough data? deep learning to the rescue!
Anaby-Tavor, A., Carmeli, B., Goldbraich, E., Kantor, A., Kour, G., Shlomov, S., Tepper, N., and Zwerdling, N. (2020) · 2020
Cited alongside, same era.
Don’t stop pretraining: Adapt language models to domains and tasks
Gururangan, S., Marasović, A., Swayamdipta, S., Lo, K., Beltagy, I., Downey, D., and Smith, N. A. (2020) · 2020
Cited alongside, same era.
Contextualizing hate speech classifiers with post-hoc explanation
Kennedy, B., Jin, X., Mostafazadeh Davani, A., Dehghani, M., and Ren, X. (2020) · 2020
Cited alongside, same era.
ALBERT: A lite BERT for self-supervised learning of language representations
Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R. (2020) · 2020
Cited alongside, same era.
Balancing via generation for multi-class text classification improvement
Tepper, N., Goldbraich, E., Zwerdling, N., Kour, G., Anaby Tavor, A., and Carmeli, B. (2020) · 2020
Cited alongside, same era.
Detecting hate speech with GPT-3
Chiu, K.-L., Collins, A., and Alexander, R. (2022) · 2022
Later among the works it cites.
ToxiGen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Hartvigsen, T., Gabriel, S., Palangi, H., Sap, M., Ray, D., and Kamar, E. (2022) · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., and Lowe, R. (2022) · 2022
Later among the works it cites.
Character-level hypernetworks for hate speech detection
Wullach, T., Adler, A., and Minkov, E. (2022) · 2022
Later among the works it cites.
Robust hate speech detection in social media: A cross-dataset empirical evaluation
Antypas, D. and Camacho-Collados, J. (2023) · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Data expansion using back translation and paraphrasing for hate speech detection
Beddiar, D. R., Jahan, M. S., and Oussalah, M. (2021) · 2021
Cited alongside, same era.
Hatexplain: A benchmark dataset for explainable hate speech detection
Mathew, B., Saha, P., Yimam, S. M., Biemann, C., Goyal, P., and Mukherjee, A. (2021) · 2021
Cited alongside, same era.
Want to reduce labeling cost? GPT-3 can help
Wang, S., Liu, Y., Xu, Y., Zhu, C., and Zeng, M. (2021) · 2021
Cited alongside, same era.
Fight fire with fire: Fine-tuning hate detectors using large samples of generated hate speech
Wullach, T., Adler, A., and Minkov, E. (2021a) · 2021
Cited alongside, same era.
Challenges in automated debiasing for toxic language detection
Zhou, X., Sap, M., Swayamdipta, S., Choi, Y., and Smith, N. A. (2021) · 2021
Cited alongside, same era.
Towards automatic generation of messages countering online hate speech and microaggressions
Ashida, M. and Komachi, M. (2022) · 2022
Cited alongside, same era.
CRUSH: Contextually regularized and user anchored self-supervised hate speech detection
Chakraborty, S., Dutta, P., Roychowdhury, S., and Mukherjee, A. (2022) · 2022
Cited alongside, same era.
Generation-based data augmentation for offensive language detection: Is it worth it?
Casula, C. and Tonelli, S. (2023) · 2023
Closest in time.
Toxicity in chatgpt: Analyzing persona-assigned language models
Deshpande, A., Murahari, V., Rajpurohit, T., Kalyan, A., and Narasimhan, K. (2023) · 2023
Closest in time.
Social world knowledge: Modeling and applications
Lotan, N. and Minkov, E. (2023) · 2023
Closest in time.
Detecting multidimensional political incivility on social media
Pendzel, S., Lotan, N., Zoizner, A., and Minkov, E. (2023) · 2023
Closest in time.
Assessing the impact of contextual information in hate speech detection
Pérez, J. M., Luque, F. M., Zayat, D., Kondratzky, M., Moro, A., Serrati, P. S., Zajac, J., Miguel, P., Debandi, N., Gravano, A., et al. (2023) · 2023
Closest in time.
A comprehensive capability analysis of gpt-3 and gpt-3.5 series models
Ye, J., Chen, X., Xu, N., Zu, C., Shao, Z., Liu, S., Cui, Y., Zhou, Z., Gong, C., Shen, Y., Zhou, J., Chen, S., Gui, T., Zhang, Q., and Huang, X. (2023) · 2023
Closest in time.