Fetching the paper…
Reading the bibliography…
Jailbreaking is an emerging adversarial attack that bypasses the safety alignment deployed in off-the-shelf large language models (LLMs) and has evolved into multiple categories: human-based, optimization-based, generation-based, and the recent indirect and multilingual jailbreaks.
SCLib: A Practical and Lightweight Defense against Component Hijacking in Android Applications
Daoyuan Wu, Yao Cheng, Debin Gao, Yingjiu Li, and Robert H. Deng · 2018
Earlier work this paper cites.
SoK: Shining light on shadow stacks
Nathan Burow, Xinping Zhang, and Mathias Payer · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
https://openai.com/index/gpt-3-apps/ , 2021
GPT-3 powers the next generation of apps · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al · 2022
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Earlier work this paper cites.
Illustrating Reinforcement Learning from Human Feedback (RLHF)
Nathan Lambert, Louis Castricato, Leandro von Werra, and Alex Havrilla · 2022
Earlier work this paper cites.
CodeGen: An open large language model for code with multi-turn program synthesis
Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong · 2022
Earlier work this paper cites.
Red teaming language models with language models
Ethan Perez, Saffron Huang, H. Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
https://github.com/verazuo/jailbreak_llms/blob/main/data/forbidden_question/forbidden_question_set_with_prompts.csv.zip , 2023
Forbidden question set with prompts · 2023
Earlier work this paper cites.
Detecting language model attacks with perplexity
Gabriel Alon and Michael Kamfonas · 2023
Earlier work this paper cites.
Abusing images and sounds for indirect instruction injection in multi-modal LLMs
Eugene Bagdasaryan, Tsung-Yin Hsieh, Ben Nassi, and Vitaly Shmatikov · 2023
Earlier work this paper cites.
Defending against alignment-breaking attacks via robustly aligned LLM
Bochuan Cao, Yuanpu Cao, Lu Lin, and Jinghui Chen · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong · 2023
Earlier work this paper cites.
Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality, March 2023
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing · 2023
Earlier work this paper cites.
Large language models for code: Security hardening and adversarial testing
Jingxuan He and Martin Vechev · 2023
Earlier work this paper cites.
A visual–language foundation model for pathology image analysis using medical Twitter
Zhi Huang, Federico Bianchi, Mert Yuksekgonul, Thomas J Montine, and James Zou · 2023
Earlier work this paper cites.
Llama Guard: LLM-based input-output safeguard for human-AI conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and Madian Khabsa · 2023
Earlier work this paper cites.
Baseline defenses for adversarial attacks against aligned language models
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping-yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein · 2023
Earlier work this paper cites.
Certifying LLM safety against adversarial prompting
Aounon Kumar, Chirag Agarwal, Suraj Srinivas, Soheil Feizi, and Hima Lakkaraju · 2023
Earlier work this paper cites.
AlpacaEval: An automatic evaluator of instruction-following models
Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Earlier work this paper cites.
CCTEST: Testing and repairing code completion systems
Zongjie Li, Chaozheng Wang, Zhibo Liu, Haoxuan Wang, Shuai Wang, and Cuiyun Gao · 2023
Earlier work this paper cites.
Prompt injection attack against LLM-integrated applications
Yi Liu, Gelei Deng, Yuekang Li, Kailong Wang, Tianwei Zhang, Yepang Liu, Haoyu Wang, Yan Zheng, and Yang Liu · 2023
Earlier work this paper cites.
Jailbreaking ChatGPT via prompt engineering: An empirical study
Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, and Yang Liu · 2023
Earlier work this paper cites.
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
Yupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia, and Neil Zhenqiang Gong · 2023
Earlier work this paper cites.
Recent advances in natural language processing via large pre-trained language models: A survey
Bonan Min, Hayley Ross, Elior Sulem, Amir Pouran Ben Veyseh, Thien Huu Nguyen, Oscar Sainz, Eneko Agirre, Ilana Heintz, and Dan Roth · 2023
Earlier work this paper cites.
GPT-4V(ision) System Card
OpenAI · 2023
Earlier work this paper cites.
LLM Self Defense: By self examination, LLMs know they are being tricked
Mansi Phute, Alec Helbling, Matthew Hull, ShengYun Peng, Sebastian Szyller, Cory Cornelius, and Duen Horng Chau · 2023
Earlier work this paper cites.
Visual adversarial examples jailbreak aligned large language models
Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Mengdi Wang, and Prateek Mittal · 2023
Earlier work this paper cites.
SmoothLLM: defending large language models against jailbreaking attacks
Alexander Robey, Eric Wong, Hamed Hassani, and George J. Pappas · 2023
Earlier work this paper cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Earlier work this paper cites.
InstructTA: Instruction-tuned targeted attack for large vision-language models
Xunguang Wang, Zhenlan Ji, Pingchuan Ma, Zongjie Li, and Shuai Wang · 2023
Earlier work this paper cites.
Jailbroken: How does LLM safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Cited alongside, same era.
Magicoder: Source code is all you need
Yuxiang Wei, Zhe Wang, Jiawei Liu, Yifeng Ding, and Lingming Zhang · 2023
Cited alongside, same era.
Jailbreak and guard aligned language models with only few in-context demonstrations
Zeming Wei, Yifei Wang, and Yisen Wang · 2023
Cited alongside, same era.
Jailbreaking GPT-4V via self-adversarial attacks with system prompts
Yuanwei Wu, Xiang Li, Yixin Liu, Pan Zhou, and Lichao Sun · 2023
Cited alongside, same era.
Defending chatgpt against jailbreak attack via self-reminders
Yueqi Xie, Jingwei Yi, Jiawei Shao, Justin Curl, Lingjuan Lyu, Qifeng Chen, Xing Xie, and Fangzhao Wu · 2023
Cited alongside, same era.
Tree of attacks: Jailbreaking black-box LLMs automatically
Anay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum Anderson, Yaron Singer, and Amin Karbasi · 2024
Closest in time.
Fight back against jailbreaking via prompt adversarial tuning
Yichuan Mo, Yuji Wang, Zeming Wei, and Yisen Wang · 2024
Closest in time.
Jailbreaking attack against multimodal large language model
Zhenxing Niu, Haodong Ren, Xinbo Gao, Gang Hua, and Rong Jin · 2024
Closest in time.
AdvPrompter: Fast adaptive adversarial prompting for LLMs
Anselm Paulus, Arman Zharmagambetov, Chuan Guo, Brandon Amos, and Yuandong Tian · 2024
Closest in time.
State-specific protein-ligand complex structure prediction with a multi-scale deep generative model
Zhuoran Qiao, Weili Nie, Arash Vahdat, Thomas F. Miller III, and Anima Anandkumar · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
LeanDojo: Theorem proving with retrieval-augmented language models
Kaiyu Yang, Aidan Swope, Alex Gu, Rahul Chalamala, Peiyang Song, Shixing Yu, Saad Godil, Ryan J Prenger, and Animashree Anandkumar · 2023
Cited alongside, same era.
Low-resource languages jailbreak GPT-4
Zheng-Xin Yong, Cristina Menghini, and Stephen H Bach · 2023
Cited alongside, same era.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen · 2023
Cited alongside, same era.
On evaluating adversarial robustness of large vision-language models
Yunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang, Chongxuan Li, Ngai-Man Man Cheung, and Min Lin · 2023
Cited alongside, same era.
Judging LLM-as-a-judge with MT-Bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Cited alongside, same era.
Large language models for information retrieval: A survey
Yutao Zhu, Huaying Yuan, Shuting Wang, Jiongnan Liu, Wenhan Liu, Chenlong Deng, Zhicheng Dou, and Ji-Rong Wen · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson · 2023
Cited alongside, same era.
Closest in time.
An empirical evaluation of LLMs for solving offensive security challenges
Minghao Shao, Boyuan Chen, Sofija Jancheska, Brendan Dolan-Gavitt, Siddharth Garg, Ramesh Karri, and Muhammad Shafique · 2024
Closest in time.
The language barrier: Dissecting safety challenges of LLMs in multilingual contexts
Lingfeng Shen, Weiting Tan, Sihao Chen, Yunmo Chen, Jingyu Zhang, Haoran Xu, Boyuan Zheng, Philipp Koehn, and Daniel Khashabi · 2024
Closest in time.
"Do Anything Now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang · 2024
Closest in time.
PAL: Proxy-guided black-box attack on large language models
Chawin Sitawarin, Norman Mu, David Wagner, and Alexandre Araujo · 2024
Closest in time.
LLM4Vuln: A unified evaluation framework for decoupling and enhancing LLMs’ vulnerability reasoning
Yuqiang Sun, Daoyuan Wu, Yue Xue, Han Liu, Wei Ma, Lyuye Zhang, Miaolei Shi, and Yang Liu · 2024
Closest in time.
GPTScan: Detecting logic vulnerabilities in smart contracts by combining GPT with program analysis
Yuqiang Sun, Daoyuan Wu, Yue Xue, Han Liu, Haijun Wang, Zhengzi Xu, Xiaofei Xie, and Yang Liu · 2024
Closest in time.
Meta llama guard 2
Llama Team · 2024
Closest in time.
Mistral-7b-instruct-v0.2
The Mistral AI Team · 2024
Closest in time.
Solving olympiad geometry without human demonstrations
Trieu H. Trinh, Yuhuai Wu, Quoc V. Le, He He, and Thang Luong · 2024
Closest in time.
LLMs can defend themselves against jailbreaking in a practical manner: A vision paper
Daoyuan Wu, Shuai Wang, Yang Liu, and Ning Liu · 2024
Closest in time.
Efficient adversarial training in LLMs with continuous attacks
Sophie Xhonneux, Alessandro Sordoni, Stephan Günnemann, Gauthier Gidel, and Leo Schwinn · 2024
Closest in time.
Fuzz4All: Universal fuzzing with large language models
Chunqiu Steven Xia, Matteo Paltenghi, Jia Le Tian, Michael Pradel, and Lingming Zhang · 2024
Closest in time.
GradSafe: Detecting unsafe prompts for LLMs via safety-critical gradient analysis
Yueqi Xie, Minghong Fang, Renjie Pi, and Neil Gong · 2024
Closest in time.
Defensive prompt patch: A robust and interpretable defense of LLMs against jailbreak attacks
Chen Xiong, Xiangyu Qi, Pin-Yu Chen, and Tsung-Yi Ho · 2024
Closest in time.
SafeDecoding: Defending against jailbreak attacks via safety-aware decoding
Zhangchen Xu, Fengqing Jiang, Luyao Niu, Jinyuan Jia, Bill Yuchen Lin, and Radha Poovendran · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan · 2024
Closest in time.
LLM-Fuzzer: Scaling assessment of large language model jailbreaks
Jiahao Yu, Xingwei Lin, Zheng Yu, and Xinyu Xing · 2024
Closest in time.
Don’t listen to me: Understanding and exploring jailbreak prompts of large language models
Zhiyuan Yu, Xiaogeng Liu, Shunning Liang, Zach Cameron, Chaowei Xiao, and Ning Zhang · 2024
Closest in time.
GPT-4 is too smart to be safe: Stealthy chat with LLMs via cipher
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu · 2024
Closest in time.
Intention analysis prompting makes large language models a good jailbreak defender
Yuqi Zhang, Liang Ding, Lefei Zhang, and Dacheng Tao · 2024
Closest in time.
Defending large language models against jailbreaking attacks through goal prioritization
Zhexin Zhang, Junxiao Yang, Pei Ke, Fei Mi, Hongning Wang, and Minlie Huang · 2024
Closest in time.
Defending large language models against jailbreak attacks via layer-specific editing
Wei Zhao, Zhe Li, Yige Li, Ye Zhang, and Jun Sun · 2024
Closest in time.
On prompt-driven safeguarding for large language models
Chujie Zheng, Fan Yin, Hao Zhou, Fandong Meng, Jie Zhou, Kai-Wei Chang, Minlie Huang, and Nanyun Peng · 2024
Closest in time.
Improved few-shot jailbreaking can circumvent aligned language models and their defenses
Xiaosen Zheng, Tianyu Pang, Chao Du, Qian Liu, Jing Jiang, and Min Lin · 2024
Closest in time.
Robust prompt optimization for defending language models against jailbreaking attacks
Andy Zhou, Bo Li, and Haohan Wang · 2024
Closest in time.
Improving alignment and robustness with circuit breakers
Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, Rowan Wang, Zico Kolter, Matt Fredrikson, and Dan Hendrycks · 2024
Closest in time.
Jailbreaking leading safety-aligned LLMs with simple adaptive attacks
Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion · 2025
Closest in time.
Improved techniques for optimization-based jailbreaking on large language models
Xiaojun Jia, Tianyu Pang, Chao Du, Yihao Huang, Jindong Gu, Yang Liu, Xiaochun Cao, and Min Lin · 2025
Closest in time.
PropertyGPT: LLM-driven formal verification of smart contracts through retrieval-augmented property generation
Ye Liu, Yue Xue, Daoyuan Wu, Yuqiang Sun, Yi Li, Miaolei Shi, and Yang Liu · 2025
Closest in time.