Fetching the paper…
Reading the bibliography…
Although Large Language Models (LLMs) have demonstrated significant capabilities in executing complex tasks in a zero-shot manner, they are susceptible to jailbreak attacks and can be manipulated to produce harmful outputs.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2018
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Gradient-based adversarial attacks against text transformers
Chuan Guo, Alexandre Sablayrolles, Hervé Jégou, and Douwe Kiela · 2021
Earlier work this paper cites.
Practical adversarial attacks on spatiotemporal traffic forecasting models
Fan Liu, Hao Liu, and Wenzhao Jiang · 2022
Earlier work this paper cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries, 2023
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong · 2023
Earlier work this paper cites.
Tree of attacks: Jailbreaking black-box llms automatically
Anay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum Anderson, Yaron Singer, and Amin Karbasi · 2023
Earlier work this paper cites.
GPTFUZZER: red teaming large language models with auto-generated jailbreak prompts
Jiahao Yu, Xingwei Lin, Zheng Yu, and Xinyu Xing · 2023
Earlier work this paper cites.
Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu · 2023
Earlier work this paper cites.
Multilingual jailbreak challenges in large language models
Yue Deng, Wenxuan Zhang, Sinno Jialin Pan, and Lidong Bing · 2023
Earlier work this paper cites.
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang · 2023
Earlier work this paper cites.
Defending chatgpt against jailbreak attack via self-reminders
Yueqi Xie, Jingwei Yi, Jiawei Shao, Justin Curl, Lingjuan Lyu, Qifeng Chen, Xing Xie, and Fangzhao Wu · 2023
Earlier work this paper cites.
Huachuan Qiu, Shuai Zhang, Anqi Li, Hongliang He, and Zhenzhong Lan · 2023
Earlier work this paper cites.
Smoothllm: Defending large language models against jailbreaking attacks
Alexander Robey, Eric Wong, Hamed Hassani, and George J. Pappas · 2023
Earlier work this paper cites.
Rain: Your language models can align themselves without finetuning
Yuhui Li, Fangyun Wei, Jinjing Zhao, Chao Zhang, and Hongyang Zhang · 2023
Earlier work this paper cites.
Large language model unlearning
Yao Yuanshun, Xu Xiaojun, and Liu Yang · 2023
Earlier work this paper cites.
Catastrophic jailbreak of open-source llms via exploiting generation
Yangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li, and Danqi Chen · 2023
Earlier work this paper cites.
Red-teaming large language models using chain of utterances for safety-alignment
Rishabh Bhardwaj and Soujanya Poria · 2023
Earlier work this paper cites.
Analyzing the inherent response tendency of llms: Real-world instructions-driven jailbreak
Yanrui Du, Sendong Zhao, Ming Ma, Yuhan Chen, and Bing Qin · 2023
Earlier work this paper cites.
Jailbreaker: Automated jailbreak across multiple large language model chatbots
Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu · 2023
Earlier work this paper cites.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Earlier work this paper cites.
Smoothllm: Defending large language models against jailbreaking attacks
Alexander Robey, Eric Wong, Hamed Hassani, and George J Pappas · 2023
Earlier work this paper cites.
Llm censorship: A machine learning challenge or a computer security problem?
David Glukhov, Ilia Shumailov, Yarin Gal, Nicolas Papernot, and Vardan Papyan · 2023
Earlier work this paper cites.
Detecting language model attacks with perplexity
Gabriel Alon and Michael Kamfonas · 2023
Earlier work this paper cites.
Defending against alignment-breaking attacks via robustly aligned llm
Bochuan Cao, Yuanpu Cao, Lu Lin, and Jinghui Chen · 2023
Earlier work this paper cites.
Federico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger, Dan Jurafsky, Tatsunori Hashimoto, and James Zou · 2023
Earlier work this paper cites.
Understanding hidden context in preference learning: Consequences for rlhf
Anand Siththaranjan, Cassidy Laidlaw, and Dylan Hadfield-Menell · 2023
Earlier work this paper cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Earlier work this paper cites.
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang · 2023
Earlier work this paper cites.
Decodingtrust: A comprehensive assessment of trustworthiness in GPT models
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, Sang T. Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li · 2023
Cited alongside, same era.
Improved baselines with visual instruction tuning
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee · 2023
Cited alongside, same era.
Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models
Erfan Shayegani, Yue Dong, and Nael Abu-Ghazaleh · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing · 2023
Cited alongside, same era.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Gpt-4 jailbreaks itself with near-perfect success using self-explanation, 2024
Govind Ramesh, Yao Dou, and Wei Xu · 2024
Closest in time.
Chain of attack: a semantic-driven contextual multi-turn attacker for llm, 2024
Xikang Yang, Xuehai Tang, Songlin Hu, and Jizhong Han · 2024
Closest in time.
Sandwich attack: Multi-language mixture adaptive attack on llms, 2024
Bibek Upadhayay and Vahid Behzadan · 2024
Closest in time.
Robust prompt optimization for defending language models against jailbreaking attacks
Andy Zhou, Bo Li, and Haohan Wang · 2024
Closest in time.
Prompt stealing attacks against large language models
Zeyang Sha and Yang Zhang · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson · 2023
Cited alongside, same era.
Autodan: Generating stealthy jailbreak prompts on aligned large language models
Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao · 2024
Cited alongside, same era.
Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi · 2024
Cited alongside, same era.
Weak-to-strong extrapolation expedites alignment
Chujie Zheng, Ziqi Wang, Heng Ji, Minlie Huang, and Nanyun Peng · 2024
Cited alongside, same era.
Jailbreaklens: Visual analysis of jailbreak attacks against large language models
Yingchaojie Feng, Zhizhang Chen, Zhining Kang, Sijia Wang, Minfeng Zhu, Wei Zhang, and Wei Chen · 2024
Cited alongside, same era.
Zeyi Liao and Huan Sun · 2024
Cited alongside, same era.
Advprompter: Fast adaptive adversarial prompting for llms
Anselm Paulus, Arman Zharmagambetov, Chuan Guo, Brandon Amos, and Yuandong Tian · 2024
Cited alongside, same era.
Tong Liu, Yingjie Zhang, Zhe Zhao, Yinpeng Dong, Guozhu Meng, and Kai Chen · 2024
Cited alongside, same era.
Jiabao Ji, Bairu Hou, Alexander Robey, George J Pappas, Hamed Hassani, Yang Zhang, Eric Wong, and Shiyu Chang · 2024
Closest in time.
Mitigating fine-tuning jailbreak attack with backdoor enhanced alignment
Jiongxiao Wang, Jiazhao Li, Yiquan Li, Xiangyu Qi, Muhao Chen, Junjie Hu, Yixuan Li, Bo Li, and Chaowei Xiao · 2024
Closest in time.
Prompt-driven llm safeguarding via directed representation optimization
Chujie Zheng, Fan Yin, Hao Zhou, Fandong Meng, Jie Zhou, Kai-Wei Chang, Minlie Huang, and Nanyun Peng · 2024
Closest in time.
Pruning for protection: Increasing jailbreak resistance in aligned llms without fine-tuning
Adib Hasan, Ileana Rugina, and Alex Wang · 2024
Closest in time.
Improving alignment and robustness with circuit breakers, 2024
Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, Rowan Wang, Zico Kolter, Matt Fredrikson, and Dan Hendrycks · 2024
Closest in time.
Eraser: Jailbreaking defense in large language models via unlearning harmful knowledge, 2024
Weikai Lu, Ziqian Zeng, Jianwei Wang, Zhengdong Lu, Zelin Chen, Huiping Zhuang, and Cen Chen · 2024
Closest in time.
The instruction hierarchy: Training llms to prioritize privileged instructions
Eric Wallace, Kai Xiao, Reimar Leike, Lilian Weng, Johannes Heidecke, and Alex Beutel · 2024
Closest in time.
Weidi Luo, Siyuan Ma, Xiaogeng Liu, Xiaoyu Guo, and Chaowei Xiao · 2024
Closest in time.
Bells: A framework towards future proof benchmarks for the evaluation of llm safeguards
Diego Dorn, Alexandre Variengien, Charbel-Raphaël Segerie, and Vincent Corruble · 2024
Closest in time.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale N Fung, and Steven Hoi · 2024
Closest in time.
Comprehensive assessment of jailbreak attacks against llms
Junjie Chu, Yugeng Liu, Ziqing Yang, Xinyue Shen, Michael Backes, and Yang Zhang · 2024
Closest in time.
Autojailbreak: Exploring jailbreak attacks and defenses through a dependency lens, 2024
Lin Lu, Hai Yan, Zenghui Yuan, Jiawen Shi, Wenqi Wei, Pin-Yu Chen, and Pan Zhou · 2024
Closest in time.
Defending llms against jailbreaking attacks via backtranslation
Yihan Wang, Zhouxing Shi, Andrew Bai, and Cho-Jui Hsieh · 2024
Closest in time.
Defensive prompt patch: A robust and interpretable defense of llms against jailbreak attacks, 2024
Chen Xiong, Xiangyu Qi, Pin-Yu Chen, and Tsung-Yi Ho · 2024
Closest in time.
Protecting your llms with information bottleneck
Zichuan Liu, Zefan Wang, Linjie Xu, Jinyu Wang, Lei Song, Tianchun Wang, Chunlin Chen, Wei Cheng, and Jiang Bian · 2024
Closest in time.
Robustifying safety-aligned large language models through clean data curation, 2024
Xiaoqun Liu, Jiacheng Liang, Muchao Ye, and Zhaohan Xi · 2024
Closest in time.
Defending large language models against jailbreak attacks via layer-specific editing, 2024
Wei Zhao, Zhe Li, Yige Li, Ye Zhang, and Jun Sun · 2024
Closest in time.
Cross-task defense: Instruction-tuning llms for content safety, 2024
Yu Fu, Wen Xiao, Jia Chen, Jiachen Li, Evangelos Papalexakis, Aichi Chien, and Yue Dong · 2024
Closest in time.
Efficient adversarial training in llms with continuous attacks, 2024
Sophie Xhonneux, Alessandro Sordoni, Stephan Günnemann, Gauthier Gidel, and Leo Schwinn · 2024
Closest in time.
Risk taxonomy, mitigation, and assessment benchmarks of large language model systems
Tianyu Cui, Yanling Wang, Chuanpu Fu, Yong Xiao, Sijia Li, Xinhao Deng, Yunpeng Liu, Qinglin Zhang, Ziyi Qiu, Peiyang Li, et al · 2024
Closest in time.
Introducing meta llama 3: The most capable openly available llm to date, 2024
Meta · 2024
Closest in time.
Adversarial tuning: Defending against jailbreak attacks for llms
Fan Liu, Zhao Xu, and Hao Liu · 2024
Closest in time.
Improved techniques for optimization-based jailbreaking on large language models, 2024
Xiaojun Jia, Tianyu Pang, Chao Du, Yihao Huang, Jindong Gu, Yang Liu, Xiaochun Cao, and Min Lin · 2024
Closest in time.