Fetching the paper…
Reading the bibliography…
Semantic cosine similarity
Faisal Rahutomo, Teruaki Kitasuka, Masayoshi Aritsugi, et al · 2012
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy · 2014
Earlier work this paper cites.
Towards evaluating the robustness of neural networks
Nicholas Carlini and David Wagner · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Adversarial examples in the physical world
Alexey Kurakin, Ian J Goodfellow, and Samy Bengio · 2018
Earlier work this paper cites.
Textbugger: Generating adversarial text against real-world applications
Jinfeng Li, Shouling Ji, Tianyu Du, Bo Li, and Ting Wang · 2018
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Ecgadv: Generating adversarial electrocardiogram to misguide arrhythmia classification system
Huangxun Chen, Chenyu Huang, Qianyi Huang, Qian Zhang, and Wei Wang · 2020
Earlier work this paper cites.
Bae: Bert-based adversarial examples for text classification
Siddhant Garg and Goutham Ramakrishnan · 2020
Earlier work this paper cites.
Nsfw words list on github, 2020
R. George · 2020
Earlier work this paper cites.
Deep learning models for electrocardiograms are susceptible to adversarial attack
Xintian Han, Yuxuan Hu, Luca Foschini, Larry Chinitz, Lior Jankelson, and Rajesh Ranganath · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
Is bert really robust? a strong baseline for natural language attack on text classification and entailment
Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits · 2020
Cited alongside, same era.
How well can text-to-image generative models understand ethical natural language interventions?
Hritik Bansal, Da Yin, Masoud Monajatipoor, and Kai-Wei Chang · 2022
Cited alongside, same era.
Nsfw text classifier on hugging face, 2022
M. Li · 2022
Cited alongside, same era.
The capacity for moral self-correction in large language models
Deep Ganguli, Amanda Askell, Nicholas Schiefer, Thomas Liao, Kamilė Lukošiūtė, Anna Chen, Anna Goldie, Azalia Mirhoseini, Catherine Olsson, Danny Hernandez, et al · 2023
Closest in time.
A survey of generative ai applications
Roberto Gozalo-Brizuela and Eduardo C Garrido-Merchán · 2023
Closest in time.
A holistic approach to undesired content detection in the real world
Todor Markov, Chong Zhang, Sandhini Agarwal, Florentine Eloundou Nekoul, Theodore Lee, Steven Adler, Angela Jiang, and Lilian Weng · 2023
Closest in time.
Adversarial prompting for black box foundation models
Natalie Maus, Patrick Chao, Eric Wong, and Jacob Gardner · 2023
Closest in time.
Prompt-specific poisoning attacks on text-to-image generative models
Shawn Shan, Wenxin Ding, Josephine Passananti, Haitao Zheng, and Ben Y. Zhao · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Raphaël Millière · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al · 2022
Cited alongside, same era.
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Cited alongside, same era.
https://www.reddit.com/r/ChatGPT/comments/11vlp7j/nsfwgpt_that_nsfw_prompt/ , 2023
Nsfw gpt · 2023
Cited alongside, same era.
Surrogateprompt: Bypassing the safety filter of text-to-image models via substitution
Zhongjie Ba, Jieming Zhong, Jiachen Lei, Peng Cheng, Qinglong Wang, Zhan Qin, Zhibo Wang, and Kui Ren · 2023
Cited alongside, same era.
https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboard
Chatbot Arena
Cited in the paper.
Closest in time.
Sneakyprompt: Evaluating robustness of text-to-image generative models’ safety filters
Yuchen Yang, Bo Hui, Haolin Yuan, Neil Gong, and Yinzhi Cao · 2023
Closest in time.
Promptbench: Towards evaluating the robustness of large language models on adversarial prompts
Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zichen Wang, Hao Chen, Yidong Wang, Linyi Yang, Wei Ye, Yue Zhang, Neil Zhenqiang Gong, et al · 2023
Closest in time.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson · 2023
Closest in time.
Coljailbreak: Collaborative generation and editing for jailbreaking text-to-image deep generation
Yizhuo Ma, Shanmin Pang, Qi Guo, Tianyu Wei, and Qing Guo · 2024
Closest in time.
Tree of attacks: Jailbreaking black-box llms automatically
Anay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum Anderson, Yaron Singer, and Amin Karbasi · 2024
Closest in time.
Trojllm: A black-box trojan prompt attack on large language models
Jiaqi Xue, Mengxin Zheng, Ting Hua, Yilin Shen, Yepeng Liu, Ladislau Bölöni, and Qian Lou · 2024
Closest in time.