Fetching the paper…
Reading the bibliography…
Text-based image generation models, such as Stable Diffusion and DALL-E 3, hold significant potential in content creation and publishing workflows, making them the focus in recent years.
Binary codes capable of correcting deletions, insertions, and reversals
Vladimir I Levenshtein and 1 others. 1966 · 1966
Earlier work this paper cites.
Bert-attack: Adversarial attack against bert using bert
Linyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue, and Xipeng Qiu. 2020 · 2004
Earlier work this paper cites.
Recipes for safety in open-domain chatbots
Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan. 2020 · 2010
Earlier work this paper cites.
The protection of children online: a brief scoping review to identify vulnerable groups
Emily R Munro. 2011 · 2011
Earlier work this paper cites.
Abusive language detection in online user content
Chikashi Nobata, Joel Tetreault, Achint Thomas, Yashar Mehdad, and Yi Chang. 2016 · 2016
Earlier work this paper cites.
Internet misconduct impact adolescent mental health in taiwan: The moderating roles of internet addiction
Tai-Kuei Yu and Cheng-Min Chao. 2016 · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aäron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. 2017 · 2017
Earlier work this paper cites.
Automatic detection of pornographic and gambling websites based on visual and textual content using a decision mechanism
Yang Chen, Rongfeng Zheng, Anmin Zhou, Shan Liao, and Liang Liu. 2020 · 2020
Earlier work this paper cites.
Seq2sick: Evaluating the robustness of sequence-to-sequence models with adversarial examples
Minhao Cheng, Jinfeng Yi, Pin-Yu Chen, Huan Zhang, and Cho-Jui Hsieh. 2020 · 2020
Earlier work this paper cites.
The hateful memes challenge: Detecting hate speech in multimodal memes
Douwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami, Amanpreet Singh, Pratik Ringshia, and Davide Testuggine. 2020 · 2020
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, and 1 others. 2022 · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, A. Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2021a · 2022
Earlier work this paper cites.
Towards adversarial attack on vision-language pre-training models
Jiaming Zhang, Qi Yi, and Jitao Sang. 2022 · 2022
Earlier work this paper cites.
The catch-22 of ai chatbots, https://www.forbes.com/sites/timbajarin/2023/09/06/the-catch-22-of-ai-chatbots/
Tim Bajarin. 2023 · 2023
Earlier work this paper cites.
Easily accessible text-to-image generation amplifies demographic stereotypes at large scale
Federico Bianchi, Pratyusha Kalluri, Esin Durmus, Faisal Ladhak, Myra Cheng, Debora Nozza, Tatsunori Hashimoto, Dan Jurafsky, James Y. Zou, and Aylin Caliskan. 2022 · 2023
Earlier work this paper cites.
Can large language models be an alternative to human evaluations?
Cheng-Han Chiang and Hung yi Lee. 2023 · 2023
Earlier work this paper cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023 · 2023
Earlier work this paper cites.
Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models
Jaemin Cho, Abhay Zala, and Mohit Bansal. 2023 · 2023
Earlier work this paper cites.
Number of midjourney users and statistics
Aiyub Dawood. 2023 · 2023
Earlier work this paper cites.
Yimo Deng and Huangxun Chen. 2023 · 2023
Cited alongside, same era.
Evaluating the robustness of text-to-image diffusion models against real-world attacks
Hongcheng Gao, Hao Zhang, Yinpeng Dong, and Zhijie Deng. 2023 · 2023
Cited alongside, same era.
Personalization as a shortcut for few-shot backdoor attack against text-to-image diffusion models
Yihao Huang, Qing Guo, and Felix Juefei-Xu. 2023 · 2023
Cited alongside, same era.
Character as pixels: A controllable prompt adversarial attacking framework for black-box text guided image generation models
Ziyi Kou, Shichao Pei, Yijun Tian, and Xiangliang Zhang. 2023 · 2023
Cited alongside, same era.
I see dead people: Gray-box adversarial attack on image-to-text models
Yilei Jiang, Weihong Li, Yiyuan Zhang, Minghong Cai, and Xiangyu Yue. 2024 · 2024
Closest in time.
Jailbreaking large language models against moderation guardrails via cipher characters
Haibo Jin, Andy Zhou, Joe D Menke, and Haohan Wang. 2024 · 2024
Closest in time.
Adversaries can misuse combinations of safe models
Erik Jones, Anca Dragan, and Jacob Steinhardt. 2024 · 2024
Closest in time.
Investigating subtler biases in llms: Ageism, beauty, institutional, and nationality bias in generative models
Mahammed Kamruzzaman, Md. Minul Islam Shovon, and Gene Louis Kim. 2024 · 2024
Closest in time.
Automatic jailbreaking of the text-to-image generative ai systems
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Raz Lapid and Moshe Sipper. 2023 · 2023
Cited alongside, same era.
Adversarial example does good: Preventing painting imitation from diffusion models via adversarial examples
Chumeng Liang, Xiaoyu Wu, Yang Hua, Jiaru Zhang, Yiming Xue, Tao Song, Zhengui Xue, Ruhui Ma, and Haibing Guan. 2023 · 2023
Cited alongside, same era.
Riatig: Reliable and imperceptible adversarial text-to-image generation with natural prompts
Han Liu, Yuhao Wu, Shixuan Zhai, Bo Yuan, and Ning Zhang. 2023 · 2023
Cited alongside, same era.
The art of deception: Black-box attack against text-to-image diffusion model
Yuetong Lu, Jingyao Xu, Yandong Li, Siyang Lu, Wei Xiang, and Wei Lu. 2023 · 2023
Cited alongside, same era.
Adversarial nibbler: A data-centric challenge for improving the safety of text-to-image models
Alicia Parrish, Hannah Rose Kirk, Jessica Quaye, Charvi Rastogi, Max Bartolo, Oana Inel, Juan Ciro, Rafael Mosquera, Addison Howard, William J. Cukierski, D. Sculley, Vijay Janapa Reddi, and Lora Aroyo. 2023 · 2023
Cited alongside, same era.
Ignore this title and hackaprompt: Exposing systemic vulnerabilities of llms through a global prompt hacking competition
Sander Schulhoff, Jeremy Pinto, Anaum Khan, Louis-Franccois Bouchard, Chenglei Si, Svetlina Anati, Valen Tagliabue, Anson Liu Kost, Christopher Carnahan, and Jordan L. Boyd-Graber. 2023 · 2023
Cited alongside, same era.
Asymmetric bias in text-to-image generation with adversarial attacks
Haz Sameen Shahgir, Xianghao Kong, Greg Ver Steeg, and Yue Dong. 2023 · 2023
Cited alongside, same era.
Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models
Erfan Shayegani, Yue Dong, and Nael Abu-Ghazaleh. 2023 · 2023
Cited alongside, same era.
Minseon Kim, Hyomin Lee, Boqing Gong, Huishuai Zhang, and Sung Ju Hwang. 2024 · 2024
Closest in time.
Weidi Luo, Siyuan Ma, Xiaogeng Liu, Xiaoyu Guo, and Chaowei Xiao. 2024 · 2024
Closest in time.
Google chief admits ‘biased’ ai tool’s photo diversity offended users
Dan Milmo and Alex Hern. 2024 · 2024
Closest in time.
Jailbreaking attack against multimodal large language model
Zhenxing Niu, Haodong Ren, Xinbo Gao, Gang Hua, and Rong Jin. 2024 · 2024
Closest in time.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. 2024 · 2024
Closest in time.
Exploring safety generalization challenges of large language models via code
Qibing Ren, Chang Gao, Jing Shao, Junchi Yan, Xin Tan, Wai Lam, and Lizhuang Ma. 2024 · 2024
Closest in time.
Nightshade: Prompt-specific poisoning attacks on text-to-image generative models
Shawn Shan, Wenxin Ding, Josephine Passananti, Haitao Zheng, and Ben Y. Zhao. 2023 · 2024
Closest in time.
New job, new gender? measuring the social bias in image generation models
Wenxuan Wang, Haonan Bai, Jen tse Huang, Yuxuan Wan, Youliang Yuan, Haoyi Qiu, Nanyun Peng, and Michael R. Lyu. 2024 · 2024
Closest in time.
Shadow alignment: The ease of subverting safely-aligned language models
Xianjun Yang, Xiao Wang, Qi Zhang, Linda Ruth Petzold, William Yang Wang, Xun Zhao, and Dahua Lin. 2024a · 2024
Closest in time.
Mma-diffusion: Multimodal attack on diffusion models
Yijun Yang, Ruiyuan Gao, Xiaosen Wang, Nan Xu, and Qiang Xu. 2023a · 2024
Closest in time.
Sneakyprompt: Jailbreaking text-to-image generative models
Yuchen Yang, Bo Hui, Haolin Yuan, Neil Gong, and Yinzhi Cao. 2023b · 2024
Closest in time.
Sneakyprompt: Jailbreaking text-to-image generative models
Yuchen Yang, Bo Hui, Haolin Yuan, Neil Gong, and Yinzhi Cao. 2024b · 2024
Closest in time.
GPT-4 is too smart to be safe: Stealthy chat with LLMs via cipher
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu. 2024 · 2024
Closest in time.
How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge AI safety by humanizing llms
Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi. 2024 · 2024
Closest in time.
Forget-me-not: Learning to forget in text-to-image diffusion models
Gong Zhang, Kai Wang, Xingqian Xu, Zhangyang Wang, and Humphrey Shi. 2024 · 2024
Closest in time.