Fetching the paper…
Reading the bibliography…
Despite remarkable advances that large language models have achieved in chatbots, maintaining a non-toxic user-AI interactive environment has become increasingly critical nowadays.
A crowd-based evaluation of abuse response strategies in conversational agents
Amanda Cercas Curry and Verena Rieser. 2019 · 1909
Earlier work this paper cites.
Toxicity detection: Does context really matter?
John Pavlopoulos, Jeffrey Sorensen, Lucas Dixon, Nithum Thain, and Ion Androutsopoulos. 2020 · 2006
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020 · 2009
Earlier work this paper cites.
Hatebert: Retraining BERT for abusive language detection in english
Tommaso Caselli, Valerio Basile, Jelena Mitrovic, and Michael Granitzer. 2020 · 2010
Earlier work this paper cites.
Fortifying toxic speech detectors against veiled toxicity
Xiaochuang Han and Yulia Tsvetkov. 2020 · 2010
Earlier work this paper cites.
Fasttext.zip: Compressing text classification models
Armand Joulin, Edouard Grave, Piotr Bojanowski, Matthijs Douze, Hérve Jégou, and Tomas Mikolov. 2016 · 2016
Earlier work this paper cites.
Hateful symbols or hateful people? predictive features for hate speech detection on twitter
Zeerak Waseem and Dirk Hovy. 2016 · 2016
Earlier work this paper cites.
# metoo alexa: how conversational systems respond to sexual harassment
Amanda Cercas Curry and Verena Rieser. 2018 · 2018
Earlier work this paper cites.
Should an agent be ignoring it? a study of verbal abuse types and conversational agents’ response styles
Hyojin Chin and Mun Yong Yi. 2019 · 2019
Earlier work this paper cites.
Exploring perceived emotional intelligence of personality-driven virtual agents in handling user challenges
Xiaojuan Ma, Emily Yang, and Pascale Fung. 2019 · 2019
Cited alongside, same era.
Unintended bias in misogyny detection
Debora Nozza, Claudia Volpetti, and Elisabetta Fersini. 2019 · 2019
Cited alongside, same era.
SemEval-2019 task 6: Identifying and categorizing offensive language in social media (OffensEval)
Marcos Zampieri, Shervin Malmasi, Preslav Nakov, Sara Rosenthal, Noura Farra, and Ritesh Kumar. 2019 · 2019
Cited alongside, same era.
Toxic, hateful, offensive or abusive? what are we really classifying? an empirical analysis of hate speech datasets
Paula Fortuna, Juan Soler, and Leo Wanner. 2020 · 2020
Cited alongside, same era.
HurtBERT: Incorporating lexical features with BERT for the detection of abusive language
Anna Koufakou, Endang Wahyu Pamungkas, Valerio Basile, and Viviana Patti. 2020 · 2020
Cited alongside, same era.
Convabuse: Data, analysis, and benchmarks for nuanced detection in conversational AI
Amanda Cercas Curry, Gavin Abercrombie, and Verena Rieser. 2021 · 2021
Later among the works it cites.
Measuring and improving model-moderator collaboration using uncertainty estimation
Ian D Kivlichan, Zi Lin, Jeremiah Liu, and Lucy Vasserman. 2021 · 2021
Later among the works it cites.
Challenges in automated debiasing for toxic language detection
Xuhui Zhou. 2021 · 2021
Later among the works it cites.
Exploring the role of grammar and word choice in bias toward african american english (aae) in hate speech classification
Camille Harris, Matan Halevy, Ayanna Howard, Amy Bruckman, and Diyi Yang. 2022 · 2022
Later among the works it cites.
Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rotten tomatoes movies and critic reviews dataset
Stefano Leone. 2020 · 2020
Cited alongside, same era.
Detect all abuse! toward universal abusive language detection models
Kunze Wang, Dong Lu, Caren Han, Siqu Long, and Josiah Poon. 2020 · 2020
Cited alongside, same era.
Differential tweetment: Mitigating racial dialect bias in harmful tweet detection
Ari Ball-Burack, Michelle Seng Ah Lee, Jennifer Cobbe, and Jatinder Singh. 2021 · 2021
Cited alongside, same era.
HateBERT: Retraining BERT for abusive language detection in English
Tommaso Caselli, Valerio Basile, Jelena Mitrović, and Michael Granitzer. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
Towards toxic positivity detection
Ishan Sanjeev Upadhyay, KV Aditya Srivatsa, and Radhika Mamidi. 2022 · 2022
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023 · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2023
Closest in time.