Fetching the paper…
Reading the bibliography…
Text-to-Image(T2I) models have achieved remarkable success in image generation and editing, yet these models still have many potential issues, particularly in generating inappropriate or Not-Safe-For-Work(NSFW) content.
Decision-based adversarial attacks: Reliable attacks against black-box machine learning models
Wieland Brendel, Jonas Rauber, and Matthias Bethge · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2017
Earlier work this paper cites.
Textbugger: Generating adversarial text against real-world applications
Jinfeng Li, Shouling Ji, Tianyu Du, Bo Li, and Ting Wang · 2018
Earlier work this paper cites.
Comdefend: An efficient image compression model to defend adversarial examples
Xiaojun Jia, Xingxing Wei, Xiaochun Cao, and Hassan Foroosh · 2019
Earlier work this paper cites.
Bae: Bert-based adversarial examples for text classification
Siddhant Garg and Goutham Ramakrishnan · 2020
Earlier work this paper cites.
Adv-watermark: A novel watermark perturbation for adversarial examples
Xiaojun Jia, Xingxing Wei, Xiaochun Cao, and Xiaoguang Han · 2020
Earlier work this paper cites.
Is bert really robust? a strong baseline for natural language attack on text classification and entailment
Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Safety Checker nested in Stable Diffusion
Safety Checker · 2021
Earlier work this paper cites.
Detoxify
Detoxify · 2022
Earlier work this paper cites.
Las-at: adversarial training with learnable attack strategy
Xiaojun Jia, Yong Zhang, Baoyuan Wu, Ke Ma, Jue Wang, and Xiaochun Cao · 2022
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Earlier work this paper cites.
Red-Teaming the Stable Diffusion Safety Filter
Javier Rando, Daniel Paleka, David Lindner, Lennart Heim, and Florian Tramèr · 2022
Earlier work this paper cites.
High-Resolution Image Synthesis with Latent Diffusion Models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Earlier work this paper cites.
Can Machines Help Us Answering Question 16 in Datasheets, and In Turn Reflecting on Inappropriate Content?
Patrick Schramowski, Christopher Tauchmann, and Kristian Kersting · 2022
Earlier work this paper cites.
LAION-5B: An Open Large-scale Dataset for Training Next Generation Image-text Models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
“real attackers don’t compute gradients”: bridging the gap between adversarial ml research and practice
Giovanni Apruzzese, Hyrum S Anderson, Savino Dambra, David Freeman, Fabio Pierazzi, and Kevin Roundy · 2023
Cited alongside, same era.
Yimo Deng and Huangxun Chen · 2023
Cited alongside, same era.
NSFW-text-classifier
NSFW-text-classifier · 2023
Later among the works it cites.
OpenAI-Moderation
OpenAI-Moderation · 2023
Later among the works it cites.
Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models
Yiting Qu, Xinyue Shen, Xinlei He, Michael Backes, Savvas Zannettou, and Yang Zhang · 2023
Later among the works it cites.
Raising the cost of malicious ai-powered image editing
Hadi Salman, Alaa Khaddaj, Guillaume Leclerc, Andrew Ilyas, and Aleksander Madry · 2023
Later among the works it cites.
Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models
Patrick Schramowski, Manuel Brack, Björn Deiseroth, and Kristian Kersting · 2023
Later among the works it cites.
Ring-A-Bell! How Reliable are Concept Removal Methods for Diffusion Models?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, and David Bau · 2023
Cited alongside, same era.
Evaluating the Robustness of Text-to-image Diffusion Models against Real-world Attacks
Hongcheng Gao, Hao Zhang, Yinpeng Dong, and Zhijie Deng · 2023
Cited alongside, same era.
A survey on transferability of adversarial examples across deep neural networks
Jindong Gu, Xiaojun Jia, Pau de Jorge, Wenqain Yu, Xinwei Liu, Avery Ma, Yuan Xun, Anjun Hu, Ashkan Khakzar, Zhijiang Li, et al · 2023
Cited alongside, same era.
Character As Pixels: A Controllable Prompt Adversarial Attacking Framework for Black-Box Text Guided Image Generation Models
Ziyi Kou, Shichao Pei, Yijun Tian, and Xiangliang Zhang · 2023
Cited alongside, same era.
Ablating Concepts in Text-to-Image Diffusion Models
Nupur Kumari, Bingliang Zhang, Sheng-Yu Wang, Eli Shechtman, Richard Zhang, and Jun-Yan Zhu · 2023
Cited alongside, same era.
Holistic Evaluation of Text-To-Image Models
Tony Lee, Michihiro Yasunaga, Chenlin Meng, Yifan Mai, Joon Sung Park, Agrim Gupta, Yunzhi Zhang, Deepak Narayanan, Hannah Benita Teufel, Marco Bellagente, Minguk Kang, Taesung Park, Jure Leskovec, Jun-Yan Zhu, Li Fei-Fei, Jiajun Wu, Stefano Ermon, and Percy Liang · 2023
Cited alongside, same era.
Adversarial example does good: Preventing painting imitation from diffusion models via adversarial examples
Chumeng Liang, Xiaoyu Wu, Yang Hua, Jiaru Zhang, Yiming Xue, Tao Song, Zhengui Xue, Ruhui Ma, and Haibing Guan · 2023
Cited alongside, same era.
Intriguing Properties of Text-guided Diffusion Models
Qihao Liu, Adam Kortylewski, Yutong Bai, Song Bai, and Alan L. Yuille · 2023
Cited alongside, same era.
Yu-Lin Tsai, Chia-Yi Hsu, Chulin Xie, Chih-Hsun Lin, Jia-You Chen, Bo Li, Pin-Yu Chen, Chia-Mu Yu, and Chun-Ying Huang · 2023
Later among the works it cites.
Sneakyprompt: Jailbreaking text-to-image generative models
Yuchen Yang, Bo Hui, Haolin Yuan, Neil Gong, and Yinzhi Cao · 2023
Later among the works it cites.
Yimeng Zhang, Jinghan Jia, Xin Chen, Aochuan Chen, Yihua Zhang, Jiancheng Liu, Ke Ding, and Sijia Liu · 2023
Later among the works it cites.
A Pilot Study of Query-Free Adversarial Attack against Stable Diffusion
Haomin Zhuang, Yihua Zhang, and Sijia Liu · 2023
Later among the works it cites.
Qi Guo, Shanmin Pang, Xiaojun Jia, and Qing Guo · 2024
Closest in time.
Safegen: Mitigating unsafe content generation in text-to-image models
Xinfeng Li, Yuchen Yang, Jiangyi Deng, Chen Yan, Yanjiao Chen, Xiaoyu Ji, and Wenyuan Xu · 2024
Closest in time.
Latent guard: a safety framework for text-to-image generation
Runtao Liu, Ashkan Khakzar, Jindong Gu, Qifeng Chen, Philip Torr, and Fabio Pizzati · 2024
Closest in time.
Jailbreaking prompt attack: A controllable adversarial attack against diffusion models
Jiachen Ma, Anda Cao, Zhiqing Xiao, Jie Zhang, Chao Ye, and Junbo Zhao · 2024
Closest in time.
Boosting transferability in vision-language attacks via diversification along the intersection region of adversarial trajectory
Sensen Gao, Xiaojun Jia, Xuhong Ren, Ivor Tsang, and Qing Guo · 2025
Closest in time.