Fetching the paper…
Reading the bibliography…
Recent work showed Best-of-N (BoN) jailbreaking using repeated use of random augmentations (such as capitalization, punctuation, etc) is effective against all major large language models (LLMs).
The use of confidence or fiducial limits illustrated in the case of the binomial
Charles J Clopper and Egon S Pearson · 1934
Earlier work this paper cites.
Pushing Boundaries or Crossing Lines? The Complex Ethics of ChatGPT Jailbreaking
Anja Boxleitner · 2023
Earlier work this paper cites.
Figstep: Jailbreaking large vision-language models via typographic visual prompts
Yichen Gong, Delong Ran, Jinyuan Liu, Conglei Wang, Tianshuo Cong, Anyu Wang, Sisi Duan, and Xiaoyun Wang · 2023
Earlier work this paper cites.
Jailbreaker: Automated jailbreak across multiple large language model chatbots
Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu · 2023
Earlier work this paper cites.
Contextual manipulations in adversarial prompting
L. Zhang et al · 2023
Earlier work this paper cites.
Defensive strategies against ai jailbreaking: Challenges and innovations
J. Li and R. Chen · 2023
Earlier work this paper cites.
Defending ChatGPT against jailbreak attack via self-reminders
Yueqi Xie, Jingwei Yi, Jiawei Shao, Justin Curl, Lingjuan Lyu, Qifeng Chen, Xing Xie, and Fangzhao Wu · 2023
Earlier work this paper cites.
Self-guard: Empower the llm to safeguard itself
Zezhong Wang, Fangkai Yang, Lu Wang, Pu Zhao, Hongru Wang, Liang Chen, Qingwei Lin, and Kam-Fai Wong · 2023
Earlier work this paper cites.
Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks
Rao Abhinav, Vashistha S., Naik Atharva, Aditya Somak, and Choudhury Monojit · 2023
Cited alongside, same era.
Ai control: Improving safety despite intentional subversion
Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, and Fabien Roger · 2023
Cited alongside, same era.
Automatic jailbreaking of the text-to-image generative ai systems
Minseon Kim, Hyomin Lee, Boqing Gong, Huishuai Zhang, and Sung Ju Hwang · 2024
Cited alongside, same era.
Haibo Jin, Leyang Hu, Xinuo Li, Peiyan Zhang, Chonghan Chen, Jun Zhuang, and Haohan Wang · 2024
Cited alongside, same era.
Don’t listen to me: Understanding and exploring jailbreak prompts of large language models
Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models
Liwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger, Faeze Brahman, Sachin Kumar, Niloofar Mireshghallah, Ximing Lu, Maarten Sap, Yejin Choi, et al · 2024
Later among the works it cites.
Jailbreakbench: An open robustness benchmark for jailbreaking large language models
Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J Pappas, Florian Tramer, et al · 2024
Later among the works it cites.
John Hughes, Sara Price, Aengus Lynch, Rylan Schaeffer, Fazl Barez, Sanmi Koyejo, Henry Sleight, Erik Jones, Ethan Perez, and Mrinank Sharma · 2024
Later among the works it cites.
Improving alignment and robustness with circuit breakers
Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, J Zico Kolter, Matt Fredrikson, and Dan Hendrycks · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhiyuan Yu, Xiaogeng Liu, Shunning Liang, Zach Cameron, Chaowei Xiao, and Ning Zhang · 2024
Cited alongside, same era.
Jailbreaking large language models with symbolic mathematics
Emet Bethany, Mazal Bethany, Juan Arturo Nolazco Flores, Sumit Kumar Jha, and Peyman Najafirad · 2024
Cited alongside, same era.
Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi · 2024
Cited alongside, same era.
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks · 2024
Later among the works it cites.
Using gpt-eliezer against chatgpt jailbreaking, 2022
Rebecca Gorman and Stuart Armstrong · 2025
Closest in time.
chatgpt-prompt-evaluator on aligned ai’s github, 2022
Rebecca Gorman and Stuart Armstrong · 2025
Closest in time.