Fetching the paper…
Reading the bibliography…
Jailbreak vulnerabilities in Large Language Models (LLMs), which exploit meticulously crafted prompts to elicit content that violates service guidelines, have captured the attention of research communities.
The art of software testing
Glenford J. Myers, Corey Sandler, and Tom Badgett, · 2012
Earlier work this paper cites.
“Evasion attacks against machine learning at test time,”
Battista Biggio, Igino Corona, Davide Maiorca, Blaine Nelson, Nedim Šrndić, Pavel Laskov, Giorgio Giacinto, and Fabio Roli, · 2013
Earlier work this paper cites.
“The limitations of deep learning in adversarial settings,”
Nicolas Papernot, Patrick McDaniel, Somesh Jha, Matt Fredrikson, Z Berkay Celik, and Ananthram Swami, · 2016
Earlier work this paper cites.
“Adversarial examples are not easily detected: Bypassing ten detection methods,”
Nicholas Carlini and David Wagner, · 2017
Earlier work this paper cites.
“The art, science, and engineering of fuzzing: A survey,”
Valentin J.M. Manès, HyungSeok Han, et al., · 2021
Earlier work this paper cites.
“Persistent anti-muslim bias in large language models,”
Abubakar Abid, Maheen Farooqi, and James Zou, · 2021
Earlier work this paper cites.
“Prompt programming for large language models: Beyond the few-shot paradigm,”
Laria Reynolds and Kyle McDonell, · 2021
Earlier work this paper cites.
“Introducing chatgpt,” https://openai.com/blog/chatgpt , 2022
OpenAI, · 2022
Earlier work this paper cites.
“Training language models to follow instructions with human feedback,”
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe, · 2022
Earlier work this paper cites.
“GPT-4 Technical Report,” 2023
OpenAI, · 2023
Earlier work this paper cites.
“Llama: Open and efficient foundation language models,” 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample, · 2023
Cited alongside, same era.
“Vicuna: An open-source chatbot impressing gpt-4 with 90% chatgpt quality - lmsys org,” https://lmsys.org/blog/2023-03-30-vicuna/
2023
Cited alongside, same era.
“How long can open-source llms truly promise on context length?,” June 2023
Dacheng Li*, Rulin Shao*, Anze Xie, Lianmin Zheng Ying Sheng, Joseph E. Gonzalez, Ion Stoica, Xuezhe Ma, and Hao Zhang, · 2023
Cited alongside, same era.
“GLM-130b: An open bilingual pre-trained model,”
Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, et al., · 2023
Cited alongside, same era.
“Tricking llms into disobedience: Understanding, analyzing, and preventing jailbreaks,” 2023
Abhinav Rao, Sachin Vashistha, Atharva Naik, Somak Aditya, and Monojit Choudhury, · 2023
Cited alongside, same era.
“Self-instruct: Aligning language models with self-generated instructions,”
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi, · 2023
Closest in time.
“Camel: Communicative agents for ”mind” exploration of large scale language model society,” 2023
Guohao Li, Hasan Abed Al Kader Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem, · 2023
Closest in time.
“Bloom: A 176b-parameter open-access multilingual language model,” 2023
BigScience Workshop: Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, et al., · 2023
Closest in time.
“Latent jailbreak: A benchmark for evaluating text safety and output robustness of large language models,” 2023
Huachuan Qiu, Shuai Zhang, et al., · 2023
Closest in time.
“Removing rlhf protections in gpt-4 via fine-tuning,” 2023
Qiusi Zhan, Richard Fang, Rohan Bindu, Akul Gupta, Tatsunori Hashimoto, and Daniel Kang, · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Jailbreaking chatgpt via prompt engineering: An empirical study,” 2023
Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, and Yang Liu, · 2023
Cited alongside, same era.
“Jailbroken: How does llm safety training fail?,” 2023
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt, · 2023
Cited alongside, same era.
“Jailbreaker: Automated jailbreak across multiple large language model chatbots,” 2023
Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu, · 2023
Cited alongside, same era.
“Multi-step jailbreaking privacy attacks on chatgpt,” 2023
Haoran Li, Dadi Guo, Wei Fan, Mingshi Xu, Jie Huang, Fanpu Meng, and Yangqiu Song, · 2023
Cited alongside, same era.
“Pretraining language models with human preferences,”
Tomasz Korbak et al., · 2023
Cited alongside, same era.
“DAN is my new friend: ChatGPT,” https://old.reddit.com/r/ChatGPT/comments/zlcyr9/dan_is_my_new_friend/
Cited in the paper.
“Jailbreak chat,” https://www.jailbreakchat.com/
Cited in the paper.
Emilio Ferrara, · 2023
Closest in time.
“A prompt pattern catalog to enhance prompt engineering with chatgpt,”
Jules White, Quchen Fu, Sam Hays, Michael Sandborn, Carlos Olea, Henry Gilbert, Ashraf Elnashar, Jesse Spencer-Smith, and Douglas C Schmidt, · 2023
Closest in time.
“Prompting ai art: An investigation into the creative skill of prompt engineering,”
Jonas Oppenlaender, Rhema Linder, and Johanna Silvennoinen, · 2023
Closest in time.
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang, · 2023
Closest in time.
“Self-deception: Reverse penetrating the semantic firewall of large language models,”
Zhenhua Wang, Wei Xie, Kai Chen, Baosheng Wang, Zhiwen Gui, and Enze Wang, · 2023
Closest in time.