Fetching the paper…
Reading the bibliography…
The rapid growth of social media platforms has raised significant concerns regarding online content toxicity.
Detecting offensive language in social media to protect adolescent online safety
Ying Chen, Yilu Zhou, Sencun Zhu, and Heng Xu. 2012 · 2012
Earlier work this paper cites.
Detecting hate speech on the world wide web
William Warner and Julia Hirschberg. 2012 · 2012
Earlier work this paper cites.
Dignity, harm, and hate speech
Robert Mark Simpson. 2013 · 2013
Earlier work this paper cites.
Hate speech detection with comment embeddings
Nemanja Djuric, Jing Zhou, Robin Morris, Mihajlo Grbovic, Vladan Radosavljevic, and Narayan Bhamidipati. 2015 · 2015
Earlier work this paper cites.
A lexicon-based approach for hate speech detection
Njagi Dennis Gitari, Zhang Zuping, Hanyurwimfura Damien, and Jun Long. 2015 · 2015
Earlier work this paper cites.
New classification models for detecting hate and violence web content
Shuhua Liu and Thomas Forss. 2015 · 2015
Earlier work this paper cites.
One-step and two-step classification for abusive language detection on twitter
Ji Ho Park and Pascale Fung. 2017 · 2017
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
‘won’t somebody please think of the children?’hate speech, harm, and childhood
Robert Mark Simpson. 2019 · 2019
Earlier work this paper cites.
Bert and fasttext embeddings for automatic detection of toxic speech
Ashwin Geet D’Sa, Irina Illina, and Dominique Fohr. 2020 · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020 · 2020
Cited alongside, same era.
TextAttack: A framework for adversarial attacks, data augmentation, and adversarial training in NLP
John Morris, Eli Lifland, Jin Yong Yoo, Jake Grigsby, Di Jin, and Yanjun Qi. 2020 · 2020
Cited alongside, same era.
HateBERT: Retraining BERT for abusive language detection in English
Tommaso Caselli, Valerio Basile, Jelena Mitrović, and Michael Granitzer. 2021 · 2021
Cited alongside, same era.
Latent hatred: A benchmark for understanding implicit hate speech
Mai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi, Jordyn Seybolt, Munmun De Choudhury, and Diyi Yang. 2021 · 2021
Cited alongside, same era.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Cited alongside, same era.
HARE: Explainable hate speech detection with step-by-step reasoning
Yongjin Yang, Joonkee Kim, Yujin Kim, Namgyu Ho, James Thorne, and Se-Young Yun. 2023 · 2023
Later among the works it cites.
Llama 3 model card
AI@Meta. 2024 · 2024
Closest in time.
Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou. 2024 · 2024
Closest in time.
From local to global: A graph rag approach to query-focused summarization
Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, and Jonathan Larson. 2024 · 2024
Closest in time.
Exploring cross-cultural differences in English hate speech annotations: From dataset construction to analysis
Nayeon Lee, Chani Jung, Junho Myung, Jiho Jin, Jose Camacho-Collados, Juho Kim, and Alice Oh. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Son T Luu and Ngan Luu-Thuy Nguyen. 2021 · 2021
Cited alongside, same era.
Hatexplain: A benchmark dataset for explainable hate speech detection
Binny Mathew, Punyajoy Saha, Seid Muhie Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee. 2021 · 2021
Cited alongside, same era.
Semeval-2021 task 5: Toxic spans detection
John Pavlopoulos, Jeffrey Sorensen, Léo Laugier, and Ion Androutsopoulos. 2021 · 2021
Cited alongside, same era.
Star: Bootstrapping reasoning with reasoning
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman. 2022 · 2022
Cited alongside, same era.
Towards building a robust toxicity predictor
Dmitriy Bespalov, Sourav Bhabesh, Yi Xiang, Liutong Zhou, and Yanjun Qi. 2023 · 2023
Cited alongside, same era.
Efficient toxic content detection by bootstrapping and distilling large language models
Jiang Zhang, Qiong Wu, Yiming Xu, Cheng Cao, Zheng Du, and Konstantinos Psounis. 2024a
Cited in the paper.
Don’t go to extremes: Revealing the excessive sensitivity and calibration limitations of LLMs in implicit hate speech detection
Min Zhang, Jianfeng He, Taoran Ji, and Chang-Tien Lu. 2024b
Cited in the paper.
Qwen2.5: A party of foundation models
Qwen Team. 2024 · 2024
Closest in time.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2024 · 2024
Closest in time.
Tree of thoughts: deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. 2024 · 2024
Closest in time.
Sft memorizes, rl generalizes: A comparative study of foundation model post-training
Tianzhe Chu, Yuexiang Zhai, Jihan Yang, Shengbang Tong, Saining Xie, Dale Schuurmans, Quoc V. Le, Sergey Levine, and Yi Ma. 2025 · 2025
Closest in time.