Fetching the paper…
Reading the bibliography…
The widespread use of social media necessitates reliable and efficient detection of offensive content to mitigate harmful effects.
Evaluation of text generation: A survey
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020 · 2006
Earlier work this paper cites.
Large scale crowdsourcing and characterization of twitter abusive behavior
Antigoni Founta, Constantinos Djouvas, Despoina Chatzakou, Ilias Leontiadis, Jeremy Blackburn, Gianluca Stringhini, Athena Vakali, Michael Sirivianos, and Nicolas Kourtellis. 2018 · 2018
Earlier work this paper cites.
Anatomy of online hate: developing a taxonomy and machine learning models for identifying and classifying hate in online news media
Joni Salminen, Hind Almerekhi, Milica Milenković, Soon-gyo Jung, Jisun An, Haewoon Kwak, and Bernard Jansen. 2018 · 2018
Earlier work this paper cites.
Semeval-2019 task 5: Multilingual detection of hate speech against immigrants and women in twitter
Valerio Basile, Cristina Bosco, Elisabetta Fersini, Debora Nozza, Viviana Patti, Francisco Manuel Rangel Pardo, Paolo Rosso, and Manuela Sanguinetti. 2019 · 2019
Earlier work this paper cites.
A benchmark dataset for learning to intervene in online hate speech
Jing Qian, Anna Bethke, Yinyin Liu, Elizabeth Belding, and William Yang Wang. 2019 · 2019
Earlier work this paper cites.
Handling imbalance issue in hate speech classification using sampling-based methods
Heng Rathpisey and Teguh Bharata Adji. 2019 · 2019
Earlier work this paper cites.
Prevalence and psychological effects of hateful speech in online college communities
Koustuv Saha, Eshwar Chandrasekharan, and Munmun De Choudhury. 2019 · 2019
Earlier work this paper cites.
Hate speech detection with machine-translated data: the role of annotation scheme, class imbalance and undersampling
Camilla Casula and Sara Tonelli. 2020 · 2020
Earlier work this paper cites.
The battle against online harmful information: The cases of fake news and hate speech
Anastasia Giachanou and Paolo Rosso. 2020 · 2020
Earlier work this paper cites.
Training question answering models from synthetic data
Raul Puri, Ryan Spring, Mohammad Shoeybi, Mostofa Patwary, and Bryan Catanzaro. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Earlier work this paper cites.
Social bias frames: Reasoning about social and power implications of language
Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A Smith, and Yejin Choi. 2020 · 2020
Earlier work this paper cites.
Directions in abusive language training data, a systematic review: Garbage in, garbage out
Bertie Vidgen and Leon Derczynski. 2020 · 2020
Earlier work this paper cites.
HateBERT: Retraining bert for abusive language detection in english
Tommaso Caselli, Valerio Basile, Mitrovic Jelena, Granitzer Michael, et al. 2021 · 2021
Earlier work this paper cites.
Detecting hate speech with GPT-3
Ke-Li Chiu, Annie Collins, and Rohan Alexander. 2021 · 2021
Earlier work this paper cites.
Latent hatred: A benchmark for understanding implicit hate speech
Mai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi, Jordyn Seybolt, Munmun De Choudhury, and Diyi Yang. 2021 · 2021
Earlier work this paper cites.
How well do hate speech, toxicity, abusive and offensive language classification models generalize across datasets?
Paula Fortuna, Juan Soler-Company, and Leo Wanner. 2021 · 2021
Earlier work this paper cites.
A survey of online hate speech through the causal lens
Antigoni Founta and Lucia Specia. 2021 · 2021
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models
Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. 2021 · 2021
Earlier work this paper cites.
Hatexplain: A benchmark dataset for explainable hate speech detection
Binny Mathew, Punyajoy Saha, Seid Muhie Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee. 2021 · 2021
Earlier work this paper cites.
MetaICL: Learning to learn in context
Sewon Min, Mike Lewis, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2021 · 2021
Earlier work this paper cites.
Data and its (dis) contents: A survey of dataset development and use in machine learning research
Amandalynne Paullada, Inioluwa Deborah Raji, Emily M Bender, Emily Denton, and Alex Hanna. 2021 · 2021
Cited alongside, same era.
Resources and benchmark corpora for hate speech detection: a systematic review
Fabio Poletto, Valerio Basile, Manuela Sanguinetti, Cristina Bosco, and Viviana Patti. 2021 · 2021
Cited alongside, same era.
HateCheck: Functional tests for hate speech detection models
Paul Röttger, Bertie Vidgen, Dong Nguyen, Zeerak Waseem, Helen Margetts, and Janet Pierrehumbert. 2021 · 2021
Cited alongside, same era.
fBERT: A neural transformer for identifying offensive content
Diptanu Sarkar, Marcos Zampieri, Tharindu Ranasinghe, and Alexander Ororbia. 2021 · 2021
Cited alongside, same era.
Introducing CAD: the contextual abuse dataset
Bertie Vidgen, Dong Nguyen, Helen Margetts, Patricia Rossini, and Rebekah Tromble. 2021a · 2021
Cited alongside, same era.
Active prompting with chain-of-thought for large language models
Shizhe Diao, Pengcheng Wang, Yong Lin, and Tong Zhang. 2023 · 2023
Later among the works it cites.
Bias runs deep: Implicit reasoning biases in persona-assigned llms
Shashank Gupta, Vaishnavi Shrivastava, Ameet Deshpande, Ashwin Kalyan, Peter Clark, Ashish Sabharwal, and Tushar Khot. 2023 · 2023
Later among the works it cites.
Large language models are reasoning teachers
Namgyu Ho, Laura Schmid, and Se-Young Yun. 2023 · 2023
Later among the works it cites.
The cot collection: Improving zero-shot and few-shot learning of language models via chain-of-thought fine-tuning
Seungone Kim, Se Joo, Doyoung Kim, Joel Jang, Seonghyeon Ye, Jamin Shin, and Minjoon Seo. 2023 · 2023
Later among the works it cites.
GPT-4 vs. GPT-3.5: A concise showdown
Anis Koubaa. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2021 · 2021
Cited alongside, same era.
Towards generalisable hate speech detection: a review on obstacles and solutions
Wenjie Yin and Arkaitz Zubiaga. 2021 · 2021
Cited alongside, same era.
Interpretable and high-performance hate and offensive speech detection
Marzieh Babaeianjelodar, Gurram Poorna Prudhvi, Stephen Lorenz, Keyu Chen, Sumona Mondal, Soumyabrata Dey, and Navin Kumar. 2022 · 2022
Cited alongside, same era.
A survey for in-context learning
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, and Zhifang Sui. 2022 · 2022
Cited alongside, same era.
Designing of prompts for hate speech recognition with in-context learning
Lawrence Han and Hao Tang. 2022 · 2022
Cited alongside, same era.
Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. 2022 · 2022
Cited alongside, same era.
An explainable AI model for hate speech detection on indonesian twitter
Muhammad Amien Ibrahim, Samsul Arifin, I Gusti Agung Anom Yudistira, Rinda Nariswari, Abdul Azis Abdillah, Nerru Pranuta Murnaka, and Puguh Wahyu Prasetyo. 2022 · 2022
Cited alongside, same era.
Bill Yuchen Lin, Abhilasha Ravichander, Ximing Lu, Nouha Dziri, Melanie Sclar, Khyathi Chandu, Chandra Bhagavatula, and Yejin Choi. 2023 · 2023
Later among the works it cites.
Towards conceptualization of “fair explanation”: Disparate impacts of anti-asian hate speech explanations on content moderators
Tin Nguyen, Jiannan Xu, Aayushi Roy, Hal Daumé III, and Marine Carpuat. 2023 · 2023
Later among the works it cites.
Probing LLMs for hate speech detection: strengths and vulnerabilities
Sarthak Roy, Ashish Harshvardhan, Animesh Mukherjee, and Punyajoy Saha. 2023 · 2023
Later among the works it cites.
Challenging big-bench tasks and whether chain-of-thought can solve them
Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc Le, Ed Chi, Denny Zhou, et al. 2023 · 2023
Later among the works it cites.
Stanford Alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Later among the works it cites.
LLM-powered data augmentation for enhanced crosslingual performance
Chenxi Whitehouse, Monojit Choudhury, and Alham Fikri Aji. 2023 · 2023
Later among the works it cites.
Hate speech is not free speech: Explainable machine learning for hate speech detection in code-mixed languages
Sargam Yadav, Abhishek Kaushik, and Kevin McDaid. 2023 · 2023
Later among the works it cites.
HARE: Explainable hate speech detection with step-by-step reasoning
Yongjin Yang, Joonkee Kim, Yujin Kim, Namgyu Ho, James Thorne, and Se-Young Yun. 2023 · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L Griffiths, Yuan Cao, and Karthik Narasimhan. 2023 · 2023
Later among the works it cites.
Offenseval 2023: Offensive language identification in the age of large language models
Marcos Zampieri, Sara Rosenthal, Preslav Nakov, Alphaeus Dmonte, and Tharindu Ranasinghe. 2023 · 2023
Later among the works it cites.
A survey of controllable text generation using transformer-based pre-trained language models
Hanqing Zhang, Haolin Song, Shaoyu Li, Ming Zhou, and Dawei Song. 2023 · 2023
Later among the works it cites.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2024 · 2024
Closest in time.
“Define Your Terms” : Enhancing efficient offensive speech classification with definition
Huy Nghiem, Umang Gupta, and Fred Morstatter. 2024 · 2024
Closest in time.
Yiqi Wang, Wentao Chen, Xiaotian Han, Xudong Lin, Haiteng Zhao, Yongfei Liu, Bohan Zhai, Jianbo Yuan, Quanzeng You, and Hongxia Yang. 2024 · 2024
Closest in time.