Fetching the paper…
Reading the bibliography…
Offensive language detection is increasingly crucial for maintaining a civilized social media platform and deploying pre-trained language models.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel. 2019 · 1909
Earlier work this paper cites.
Does gender matter? towards fairness in dialogue systems
Haochen Liu, Jamell Dacon, Wenqi Fan, Hui Liu, Zitao Liu, and Jiliang Tang. 2019 · 1910
Earlier work this paper cites.
Measuring nominal scale agreement among many raters
Joseph L Fleiss. 1971 · 1971
Earlier work this paper cites.
Crows-pairs: A challenge dataset for measuring social biases in masked language models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R Bowman. 2020 · 2010
Earlier work this paper cites.
Cross language text classification by model translation and semi-supervised learning
Lei Shi, Rada Mihalcea, and Mingjun Tian. 2010 · 2010
Earlier work this paper cites.
Detecting hate speech on the world wide web
William Warner and Julia Hirschberg. 2012 · 2012
Earlier work this paper cites.
Automated hate speech detection and the problem of offensive language
Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017 · 2017
Earlier work this paper cites.
A survey on hate speech detection using natural language processing
Anna Schmidt and Michael Wiegand. 2017 · 2017
Earlier work this paper cites.
Rephrasing profanity in Chinese text
Hui-Po Su, Zhen-Jie Huang, Hao-Tsung Chang, and Chuan-Jie Lin. 2017 · 2017
Earlier work this paper cites.
Understanding abuse: A typology of abusive language detection subtasks
Zeerak Waseem, Thomas Davidson, Dana Warmsley, and Ingmar Weber. 2017 · 2017
Earlier work this paper cites.
Ex machina: Personal attacks seen at scale
Ellery Wulczyn, Nithum Thain, and Lucas Dixon. 2017 · 2017
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M Bender and Batya Friedman. 2018 · 2018
Earlier work this paper cites.
Toxic comment classification challenge
Jigsaw. 2018 · 2018
Earlier work this paper cites.
Machine learning suites for online toxicity detection
David Noever. 2018 · 2018
Earlier work this paper cites.
Towards identifying social bias in dialog systems: Frame, datasets, and benchmarks
Jingyan Zhou, Jiawen Deng, Fei Mi, Yitong Li, Yasheng Wang, Minlie Huang, Xin Jiang, Qun Liu, and Helen Meng. 2022 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Build it break it fix it for dialogue safety: Robustness from adversarial human attack
Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston. 2019 · 2019
Earlier work this paper cites.
Hate speech detection: Challenges and solutions
Sean MacAvaney, Hao-Ren Yao, Eugene Yang, Katina Russell, Nazli Goharian, and Ophir Frieder. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Cited alongside, same era.
Mc-bert4hate: Hate speech detection using multi-channel bert for different languages and translations
Hajung Sohn and Hyunju Lee. 2019 · 2019
Cited alongside, same era.
Predicting the type and target of offensive posts in social media
Marcos Zampieri, Shervin Malmasi, Preslav Nakov, Sara Rosenthal, Noura Farra, and Ritesh Kumar. 2019 · 2019
Cited alongside, same era.
RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020 · 2020
Cited alongside, same era.
Fortifying Toxic Speech Detectors Against Veiled Toxicity
Anticipating safety issues in e2e conversational ai: Framework and tooling
Emily Dinan, Gavin Abercrombie, A. Stevie Bergman, Shannon Spruit, Dirk Hovy, Y-Lan Boureau, and Verena Rieser. 2021 · 2021
Later among the works it cites.
A systematic review of hate speech automatic detection using natural language processing
Md Saroar Jahan and Mourad Oussalah. 2021 · 2021
Later among the works it cites.
SWSR: A Chinese dataset and lexicon for online sexism detection
Aiqi Jiang, Xiaohan Yang, Yang Liu, and Arkaitz Zubiaga. 2022 · 2021
Later among the works it cites.
Capturing covertly toxic speech via crowdsourcing
Alyssa Lees, Daniel Borkan, Ian Kivlichan, Jorge Nario, and Tesh Goyal. 2021 · 2021
Later among the works it cites.
On-the-fly controlled text generation with experts and anti-experts
Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A Smith, and Yejin Choi. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xiaochuang Han and Yulia Tsvetkov. 2020 · 2020
Cited alongside, same era.
Detoxify
Laura Hanu and Unitary team. 2020 · 2020
Cited alongside, same era.
Recipes for building an open-domain chatbot
Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Kurt Shuster, Eric M. Smith, Y-Lan Boureau, and Jason Weston. 2020 · 2020
Cited alongside, same era.
Categorizing Offensive Language in Social Networks: A Chinese Corpus, Systems and an Explanation Tool
Xiangru Tang, Xianjun Shen, Yujie Wang, and Yujuan Yang. 2020 · 2020
Cited alongside, same era.
Directions in abusive language training data, a systematic review: Garbage in, garbage out
Bertie Vidgen and Leon Derczynski. 2020 · 2020
Cited alongside, same era.
A large-scale chinese short-text conversation dataset
Yida Wang, Pei Ke, Yinhe Zheng, Kaili Huang, Yong Jiang, Xiaoyan Zhu, and Minlie Huang. 2020 · 2020
Cited alongside, same era.
Recipes for safety in open-domain chatbots
Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan. 2020 · 2020
Cited alongside, same era.
Exploring stylometric and emotion-based features for multilingual cross-domain hate speech detection
Ilia Markov, Nikola Ljubešić, Darja Fišer, and Walter Daelemans. 2021 · 2021
Later among the works it cites.
StereoSet: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy. 2020 · 2021
Later among the works it cites.
Exposing the limits of zero-shot cross-lingual hate speech detection
Debora Nozza. 2021 · 2021
Later among the works it cites.
Probing toxic content in large pre-trained language models
Nedjma Djouhra Ousidhoum, Xinran Zhao, Tianqing Fang, Yangqiu Song, and Dit Yan Yeung. 2021 · 2021
Later among the works it cites.
Few-shot instruction prompts for pretrained language models to detect social biases
Shrimai Prabhumoye, Rafal Kocielnik, Mohammad Shoeybi, Anima Anandkumar, and Bryan Catanzaro. 2021 · 2021
Later among the works it cites.
Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp
Timo Schick, Sahana Udupa, and Hinrich Schütze. 2021 · 2021
Later among the works it cites.
Language models have a moral dimension
Patrick Schramowski, Cigdem Turan, Nico Andersen, Constantin Rothkopf, and Kristian Kersting. 2021 · 2021
Later among the works it cites.
"nice try, kiddo": Investigating ad hominems in dialogue responses
Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. 2021 · 2021
Later among the works it cites.
On the safety of conversational models: Taxonomy, dataset, and benchmark
Hao Sun, Guangxuan Xu, Jiawen Deng, Jiale Cheng, Chujie Zheng, Hao Zhou, Nanyun Peng, Xiaoyan Zhu, and Minlie Huang. 2021 · 2021
Later among the works it cites.
Bot-adversarial dialogue for safe conversational agents
Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan. 2021 · 2021
Later among the works it cites.
Eva: An open-domain chinese dialogue system with large-scale generative pre-training
Hao Zhou, Pei Ke, Zheng Zhang, Yuxian Gu, Yinhe Zheng, Chujie Zheng, Yida Wang, Chen Henry Wu, Hao Sun, Xiaocong Yang, et al. 2021 · 2021
Later among the works it cites.