Fetching the paper…
Reading the bibliography…
With the advent of text-to-image models and concerns about their misuse, developers are increasingly relying on image safety classifiers to moderate their generated unsafe images.
Measuring Nominal Scale Agreement Among Many Raters
Joseph L Fleiss · 1971
Earlier work this paper cites.
Fleiss’ kappa statistic without paradoxes
Rosa Falotico and Piero Quatto · 2015
Earlier work this paper cites.
Explaining and Harnessing Adversarial Examples
Ian Goodfellow, Jonathon Shlens, and Christian Szegedy · 2015
Earlier work this paper cites.
Deepfool: A Simple and Accurate Method to Fool Deep Neural Networks
Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, and Pascal Frossard · 2016
Earlier work this paper cites.
Universal Adversarial Perturbations
Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, Omar Fawzi, and Pascal Frossard · 2017
Earlier work this paper cites.
Defining Digital Self-Harm
Jessica Pater and Elizabeth D. Mynatt · 2017
Earlier work this paper cites.
Protest Activity Detection and Perceived Violence Estimation from Social Media Images
Donghyeon Won, Zachary C. Steinert-Threlkeld, and Jungseock Joo · 2017
Earlier work this paper cites.
The Socio-Moral Image Database (SMID): A novel stimulus set for the study of social, moral and affective processes
Damien L. Crone, Stefan Bode, Carsten Murawski, and Simon M. Laham · 2018
Earlier work this paper cites.
Towards Deep Learning Models Resistant to Adversarial Attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2018
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Hate Speech in Pixels: Detection of Offensive Memes towards Automatic Moderation
Benet Oriol Sabat, Cristian Canton-Ferrer, and Xavier Giró-i-Nieto · 2019
Earlier work this paper cites.
The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
Douwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami, Amanpreet Singh, Pratik Ringshia, and Davide Testuggine · 2020
Earlier work this paper cites.
A Quantitative Approach to Understanding Online Antisemitism
Savvas Zannettou, Joel Finkelstein, Barry Bradlyn, and Jeremy Blackburn · 2020
Earlier work this paper cites.
Multimodal Datasets: Misogyny, Pornography, and Malignant Stereotypes
Abeba Birhane, Vinay Uday Prabhu, and Emmanuel Kahembwe · 2021
Earlier work this paper cites.
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Earlier work this paper cites.
LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki · 2021
Earlier work this paper cites.
OpenPrompt: An Open-source Framework for Prompt-learning
Ning Ding, Shengding Hu, Weilin Zhao, Yulin Chen, Zhiyuan Liu, Haitao Zheng, and Maosong Sun · 2022
Cited alongside, same era.
Understanding and Detecting Hateful Content using Contrastive Learning
Felipe González-Pizarro and Savvas Zannettou · 2022
Cited alongside, same era.
Optimizing Prompts for Text-to-Image Generation
Yaru Hao, Zewen Chi, Li Dong, and Furu Wei · 2022
Cited alongside, same era.
LoRA: Low-Rank Adaptation of Large Language Models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Cited alongside, same era.
Hatemoji: A Test Suite and Adversarially-Generated Dataset for Benchmarking and Detecting Emoji-Based Hate
Hannah Kirk, Bertie Vidgen, Paul Röttger, Tristan Thrush, and Scott A. Hale · 2022
Cited alongside, same era.
Erasing Concepts from Diffusion Models
Rohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, and David Bau · 2023
Later among the works it cites.
Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and Madian Khabsa · 2023
Later among the works it cites.
Self-Discovering Interpretable Diffusion Latent Directions for Responsible Text-to-Image Generation
Hang Li, Chengzhi Shen, Philip H. S. Torr, Volker Tresp, and Jindong Gu · 2023
Later among the works it cites.
Visual Instruction Tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Red-Teaming the Stable Diffusion Safety Filter
Javier Rando, Daniel Paleka, David Lindner, Lennart Heim, and Florian Tramèr · 2022
Cited alongside, same era.
High-Resolution Image Synthesis with Latent Diffusion Models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Cited alongside, same era.
Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models
Patrick Schramowski, Manuel Brack, Björn Deiseroth, and Kristian Kersting · 2022
Cited alongside, same era.
Can Machines Help Us Answering Question 16 in Datasheets, and In Turn Reflecting on Inappropriate Content?
Patrick Schramowski, Christopher Tauchmann, and Kristian Kersting · 2022
Cited alongside, same era.
LAION-5B: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev · 2022
Cited alongside, same era.
https://gnet-research.org/2023/11/13/for-the-lulz-ai-generated-subliminal-hate-is-a-new-challenge-in-the-fight-against-online-harm/
AI-Generated Unsafe Image · 2023
Cited alongside, same era.
https://lmsys.org/blog/2023-03-30-vicuna/
Vicuna · 2023
Cited alongside, same era.
MaNinareh Mehrabi, Palash Goyal, Christophe Dupuy, Qian Hu, Shalini Ghosh, Richard S. Zemel, Kai-Wei Chang, Aram Galstyan, and Rahul Gupta · 2023
Later among the works it cites.
On the Evolution of (Hateful) Memes by Means of Multimodal Contrastive Learning
Yiting Qu, Xinlei He, Shannon Pierson, Michael Backes, Yang Zhang, and Savvas Zannettou · 2023
Later among the works it cites.
Unsafe Diffusion: On the Generation of Unsafe Images and Hateful Memes From Text-To-Image Models
Yiting Qu, Xinyue Shen, Xinlei He, Michael Backes, Savvas Zannettou, and Yang Zhang · 2023
Later among the works it cites.
LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Later among the works it cites.
Ring-A-Bell! How Reliable are Concept Removal Methods for Diffusion Models?
Yu-Lin Tsai, Chia-Yi Hsu, Chulin Xie, Chih-Hsun Lin, Jia-You Chen, Bo Li, Pin-Yu Chen, Chia-Mu Yu, and Chun-Ying Huang · 2023
Later among the works it cites.
Stereotypes and Smut: The (Mis)representation of Non-cisgender Identities by Text-to-Image Models
Eddie L. Ungless, Björn Ross, and Anne Lauscher · 2023
Later among the works it cites.
On the Proactive Generation of Unsafe Images From Text-To-Image Models Using Benign Prompts
Yixin Wu, Ning Yu, Michael Backes, Yun Shen, and Yang Zhang · 2023
Later among the works it cites.
SneakyPrompt: Evaluating Robustness of Text-to-image Generative Models’ Safety Filters
Yuchen Yang, Bo Hui, Haolin Yuan, Neil Zhenqiang Gong, and Yinzhi Cao · 2023
Later among the works it cites.
Yimeng Zhang, Jinghan Jia, Xin Chen, Aochuan Chen, Yihua Zhang, Jiancheng Liu, Ke Ding, and Sijia Liu · 2023
Later among the works it cites.
Moderating Illicit Online Image Promotion for Unsafe User-Generated Content Games Using Large Vision-Language Models
Keyan Guo, Ayush Utkarsh, Wenbo Ding, Isabelle Ondracek, Ziming Zhao, Guo Freeman, Nishant Vishwamitra, and Hongxin Hu · 2024
Closest in time.
LLavaGuard: VLM-based Safeguards for Vision Dataset Curation and Safety Assessment
Lukas Helff, Felix Friedrich, Manuel Brack, Kristian Kersting, and Patrick Schramowski · 2024
Closest in time.
Zero shot VLMs for hate meme detection: Are we there yet?
Naquee Rizwan, Paramananda Bhaskar, Mithun Das, Swadhin Satyaprakash Majhi, Punyajoy Saha, and Animesh Mukherjee · 2024
Closest in time.