Fetching the paper…
Reading the bibliography…
Hate speech is a harmful form of online expression, often manifesting as derogatory posts.
Quantifying the carbon emissions of machine learning
Alexandre Lacoste, Alexandra Luccioni, Victor Schmidt, and Thomas Dandres. 2019 · 1910
Earlier work this paper cites.
A coefficient of agreement for nominal scales
Jacob Cohen. 1960 · 1960
Earlier work this paper cites.
Countering online hate speech
Iginio Gagliardone, Danit Gal, Thiago Alves, and Gabriela Martinez. 2015 · 2015
Earlier work this paper cites.
Counterspeech on twitter: A field study
Susan Benesch, Derek Ruths, Kelly P Dillon, Haji Mohammad Saleem, and Lucas Wright. 2016 · 2016
Earlier work this paper cites.
Hateful symbols or hateful people? predictive features for hate speech detection on Twitter
Zeerak Waseem and Dirk Hovy. 2016 · 2016
Earlier work this paper cites.
Mean birds: Detecting aggression and bullying on twitter
Despoina Chatzakou, Nicolas Kourtellis, Jeremy Blackburn, Emiliano De Cristofaro, Gianluca Stringhini, and Athena Vakali. 2017 · 2017
Earlier work this paper cites.
Automated hate speech detection and the problem of offensive language
Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017 · 2017
Earlier work this paper cites.
Protecting chatbots from toxic content
Guillaume Baudart, Julian Dolby, Evelyn Duesterwald, Martin Hirzel, and Avraham Shinnar. 2018 · 2018
Earlier work this paper cites.
A socio-contextual approach in automated detection of public cyberbullying on twitter
Nargess Tahmasbi and Elham Rastegari. 2018 · 2018
Earlier work this paper cites.
Finding microaggressions in the wild: A case for locating elusive phenomena in social media posts
Luke Breitfeller, Emily Ahn, David Jurgens, and Yulia Tsvetkov. 2019 · 2019
Earlier work this paper cites.
Should an agent be ignoring it? a study of verbal abuse types and conversational agents’ response styles
Hyojin Chin and Mun Yong Yi. 2019 · 2019
Earlier work this paper cites.
CONAN - COunter NArratives through nichesourcing: a multilingual dataset of responses to fight online hate speech
Yi-Ling Chung, Elizaveta Kuzmenko, Serra Sinem Tekiroglu, and Marco Guerini. 2019 · 2019
Earlier work this paper cites.
A just and comprehensive strategy for using NLP to address online abuse
David Jurgens, Libby Hemphill, and Eshwar Chandrasekharan. 2019 · 2019
Earlier work this paper cites.
Thou shalt not hate: Countering online hate speech
Binny Mathew, Punyajoy Saha, Hardik Tharad, Subham Rajgaria, Prajwal Singhania, Suman Kalyan Maity, Pawan Goyal, and Animesh Mukherjee. 2019 · 2019
Earlier work this paper cites.
A benchmark dataset for learning to intervene in online hate speech
Jing Qian, Anna Bethke, Yinyin Liu, Elizabeth Belding, and William Yang Wang. 2019 · 2019
Earlier work this paper cites.
Empathy is all you need: How a conversational agent should respond to verbal abuse
Hyojin Chin, Lebogang Wame Molefi, and Mun Yong Yi. 2020 · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020 · 2020
Earlier work this paper cites.
XHate-999: Analyzing and detecting abusive language across domains and languages
Goran Glavaš, Mladen Karan, and Ivan Vulić. 2020 · 2020
Earlier work this paper cites.
Hate towards the political opponent: A Twitter corpus study of the 2020 US elections on the basis of offensive speech and stance detection
Lara Grimminger and Roman Klinger. 2021 · 2020
Earlier work this paper cites.
Social bias frames: Reasoning about social and power implications of language
Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. 2020 · 2020
Earlier work this paper cites.
Generating counter narratives against online hate speech: Data and strategies
Serra Sinem Tekiroğlu, Yi-Ling Chung, and Marco Guerini. 2020 · 2020
Earlier work this paper cites.
DIALOGPT : Large-scale generative pre-training for conversational response generation
Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. 2020 · 2020
Earlier work this paper cites.
The design and implementation of xiaoice, an empathetic social chatbot
Li Zhou, Jianfeng Gao, Di Li, and Heung-Yeung Shum. 2020 · 2020
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Earlier work this paper cites.
HateBERT: Retraining BERT for abusive language detection in English
Tommaso Caselli, Valerio Basile, Jelena Mitrović, and Michael Granitzer. 2021 · 2021
Earlier work this paper cites.
ConvAbuse: Data, analysis, and benchmarks for nuanced abuse detection in conversational AI
Amanda Cercas Curry, Gavin Abercrombie, and Verena Rieser. 2021 · 2021
Cited alongside, same era.
Towards knowledge-grounded counter narrative generation for hate speech
Yi-Ling Chung, Serra Sinem Tekiroğlu, and Marco Guerini. 2021 · 2021
Cited alongside, same era.
Latent hatred: A benchmark for understanding implicit hate speech
Mai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi, Jordyn Seybolt, Munmun De Choudhury, and Diyi Yang. 2021 · 2021
Cited alongside, same era.
Human-in-the-loop for data collection: a multi-target counter narrative dataset to fight online hate speech
Margherita Fanton, Helena Bonaldi, Serra Sinem Tekiroğlu, and Marco Guerini. 2021 · 2021
Cited alongside, same era.
Multimodal hate speech detection in greek social media
Konstantinos Perifanos and Dionysis Goutsos. 2021 · 2021
Cited alongside, same era.
OpenAI, :, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, …, and Barret Zoph. 2023 · 2023
Later among the works it cites.
On the risk of misinformation pollution with large language models
Yikang Pan, Liangming Pan, Wenhu Chen, Preslav Nakov, Min-Yen Kan, and William Wang. 2023 · 2023
Later among the works it cites.
Respectful or toxic? using zero-shot learning with language models to detect hate speech
Flor Miriam Plaza-del arco, Debora Nozza, and Dirk Hovy. 2023 · 2023
Later among the works it cites.
Probing LLMs for hate speech detection: strengths and vulnerabilities
Sarthak Roy, Ashish Harshvardhan, Animesh Mukherjee, and Punyajoy Saha. 2023 · 2023
Later among the works it cites.
People make better edits: Measuring the efficacy of LLM-generated counterfactually augmented data for harmful language detection
Indira Sen, Dennis Assenmacher, Mattia Samory, Isabelle Augenstein, Wil Aalst, and Claudia Wagner. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Eric Michael Smith, Y-Lan Boureau, and Jason Weston. 2021 · 2021
Cited alongside, same era.
Learning from the worst: Dynamically generated datasets to improve online hate detection
Bertie Vidgen, Tristan Thrush, Zeerak Waseem, and Douwe Kiela. 2021 · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan. 2022 · 2022
Cited alongside, same era.
Human-machine collaboration approaches to build a dialogue dataset for hate speech countering
Helena Bonaldi, Sara Dellantonio, Serra Sinem Tekiroğlu, and Marco Guerini. 2022 · 2022
Cited alongside, same era.
On the origin of hallucinations in conversational models: Is it the datasets or the models?
Nouha Dziri, Sivan Milton, Mo Yu, Osmar Zaiane, and Siva Reddy. 2022 · 2022
Cited alongside, same era.
ToxiGen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. 2022 · 2022
Cited alongside, same era.
Generalizable implicit hate speech detection using contrastive learning
Youngwook Kim, Shinwoo Park, and Yo-Sub Han. 2022 · 2022
Cited alongside, same era.
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, …, and Thomas Scialom. 2023 · 2023
Later among the works it cites.
Self-guard: Empower the llm to safeguard itself
Zezhong Wang, Fangkai Yang, Lu Wang, Pu Zhao, Hongru Wang, Liang Chen, Qingwei Lin, and Kam-Fai Wong. 2023 · 2023
Later among the works it cites.
Multi-party chat: Conversational agents in group settings with humans and models
Jimmy Wei, Kurt Shuster, Arthur Szlam, Jason Weston, Jack Urbanek, and Mojtaba Komeili. 2023 · 2023
Later among the works it cites.
A comprehensive capability analysis of gpt-3 and gpt-3.5 series models
Junjie Ye, Xuanting Chen, Nuo Xu, Can Zu, Zekai Shao, Shichun Liu, …, and Xuanjing Huang. 2023 · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2023
Later among the works it cites.
Multi-party multimodal conversations between patients, their companions, and a social robot in a hospital memory clinic
Angus Addlesee, Neeraj Cherakara, Nivan Nelson, Daniel Hernandez Garcia, Nancie Gunson, Weronika Sieińska, Christian Dondrup, and Oliver Lemon. 2024 · 2024
Closest in time.
Large language models are vulnerable to bait-and-switch attacks for generating harmful content
Federico Bianchi and James Zou. 2024 · 2024
Closest in time.
Outcome-constrained large language models for countering hate speech
Lingzi Hong, Pengcheng Luo, Eduardo Blanco, and Xiaoying Song. 2024 · 2024
Closest in time.
From bytes to borsch: Fine-tuning gemma and mistral for the Ukrainian language representation
Artur Kiulian, Anton Polishko, Mykola Khandoga, Oryna Chubych, Jack Connor, Raghav Ravishankar, and Adarsh Shirawalmath. 2024 · 2024
Closest in time.
Online hate and harassment: The american experience 2024
Anti-Defamation League. 2024 · 2024
Closest in time.
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, …, and Barret Zoph. 2024 · 2024
Closest in time.
Misconception
n.d. Oxford English Dictionary. 2024 · 2024
Closest in time.
Metahate: A dataset for unifying efforts on hate speech detection
Paloma Piot, Patricia Martín-Rodilla, and Javier Parapar. 2024 · 2024
Closest in time.
Llms among us: Generative ai participating in digital discourse
Kristina Radivojevic, Nicholas Clark, and Paul Brenner. 2024 · 2024
Closest in time.
Paul Röttger, Fabio Pernisi, Bertie Vidgen, and Dirk Hovy. 2024 · 2024
Closest in time.
Detecting offensive language in an open chatbot platform
Hyeonho Song, Jisu Hong, Chani Jung, Hyojin Chin, Mingi Shin, Yubin Choi, Junghoi Choi, and Meeyoung Cha. 2024 · 2024
Closest in time.
The state of online harassment
Emily A. Vogels. 2021 · 2024
Closest in time.
Zheyang Xiong, Vasilis Papageorgiou, Kangwook Lee, and Dimitris Papailiopoulos. 2024 · 2024
Closest in time.
Hate cannot drive out hate: Forecasting conversation incivility following replies to hate speech
Xinchen Yu, Eduardo Blanco, and Lingzi Hong. 2024 · 2024
Closest in time.