Fetching the paper…
Reading the bibliography…
Jailbreak attacks induce Large Language Models (LLMs) to generate harmful responses, posing severe misuse threats.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou · 2022
Earlier work this paper cites.
Red-teaming large language models using chain of utterances for safety-alignment, 2023
Rishabh Bhardwaj and Soujanya Poria · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries, 2023
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong · 2023
Earlier work this paper cites.
Safe rlhf: Safe reinforcement learning from human feedback
Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang, and Yaodong Yang · 2023
Earlier work this paper cites.
Attack prompt generation for red teaming and defending large language models, 2023
Boyi Deng, Wenjie Wang, Fuli Feng, Yang Deng, Qifan Wang, and Xiangnan He · 2023
Earlier work this paper cites.
Decoding the threat landscape: Chatgpt, fraudgpt, and wormgpt in social engineering attacks
Polra Victor Falade · 2023
Earlier work this paper cites.
Mart: Improving llm safety with multi-round automatic red-teaming, 2023
Suyu Ge, Chunting Zhou, Rui Hou, Madian Khabsa, Yi-Chia Wang, Qifan Wang, Jiawei Han, and Yuning Mao · 2023
Earlier work this paper cites.
Figstep: Jailbreaking large vision-language models via typographic visual prompts, 2023
Yichen Gong, Delong Ran, Jinyuan Liu, Conglei Wang, Tianshuo Cong, Anyu Wang, Sisi Duan, and Xiaoyun Wang · 2023
Earlier work this paper cites.
Intelligent virtual assistants with llm-based process automation, 2023
Yanchu Guan, Dong Wang, Zhixuan Chu, Shiyu Wang, Feiyue Ni, Ruihua Song, Longfei Li, Jinjie Gu, and Chenyi Zhuang · 2023
Earlier work this paper cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations, 2023
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and Madian Khabsa · 2023
Earlier work this paper cites.
Baseline defenses for adversarial attacks against aligned language models, 2023
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein · 2023
Earlier work this paper cites.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Jiaming Ji, Mickel Liu, Josef Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Chen, Ruiyang Sun, Yizhou Wang, and Yaodong Yang · 2023
Earlier work this paper cites.
Ai alignment: A comprehensive survey, 2023
Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Jiayi Zhou, Zhaowei Zhang, Fanzhi Zeng, Kwan Yee Ng, Juntao Dai, Xuehai Pan, Aidan O’Gara, Yingshan Lei, Hua Xu, Brian Tse, Jie Fu, Stephen McAleer, Yaodong Yang, Yizhou Wang, Song-Chun Zhu, Yike Guo, and Wen Gao · 2023
Earlier work this paper cites.
Open sesame! universal black box jailbreaking of large language models, 2023
Raz Lapid, Ron Langberg, and Moshe Sipper · 2023
Earlier work this paper cites.
Robustness over time: Understanding adversarial examples’ effectiveness on longitudinal versions of large language models, 2023
Yugeng Liu, Tianshuo Cong, Zhengyu Zhao, Michael Backes, Yun Shen, and Yang Zhang · 2023
Earlier work this paper cites.
An attacker’s dream? exploring the capabilities of chatgpt for developing malware
Yin Minn Pa Pa, Shunsuke Tanizaki, Tetsui Kou, Michel van Eeten, Katsunari Yoshioka, and Tsutomu Matsumoto · 2023
Earlier work this paper cites.
Latent jailbreak: A benchmark for evaluating text safety and output robustness of large language models, 2023
Huachuan Qiu, Shuai Zhang, Anqi Li, Hongliang He, and Zhenzhong Lan · 2023
Earlier work this paper cites.
Scalable and transferable black-box jailbreaks for language models via persona modulation, 2023
Rusheb Shah, Quentin Feuillade-Montixi, Soroush Pour, Arush Tagade, Stephen Casper, and Javier Rando · 2023
Earlier work this paper cites.
“do anything now”: Characterizing and evaluating in-the-wild jailbreak prompts on large language models, 2023
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang · 2023
Earlier work this paper cites.
Comparing traditional and llm-based search for consumer choice: A randomized experiment, 2023
Sofia Eleni Spatharioti, David M. Rothschild, Daniel G. Goldstein, and Jake M. Hofman · 2023
Earlier work this paper cites.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Earlier work this paper cites.
Jailbreak and guard aligned language models with only few in-context demonstrations, 2023
Zeming Wei, Yifei Wang, and Yisen Wang · 2023
Earlier work this paper cites.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts, 2023
Jiahao Yu, Xingwei Lin, Zheng Yu, and Xinyu Xing · 2023
Earlier work this paper cites.
Make them spill the beans! coercive knowledge extraction from (production) llms, 2023
Zhuo Zhang, Guangyu Shen, Guanhong Tao, Siyuan Cheng, and Xiangyu Zhang · 2023
Earlier work this paper cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E Gonzalez, and Ion Stoica · 2023
Earlier work this paper cites.
Autodan: Interpretable gradient-based adversarial attacks on large language models, 2023
Sicheng Zhu, Ruiyi Zhang, Bang An, Gang Wu, Joe Barrow, Zichao Wang, Furong Huang, Ani Nenkova, and Tong Sun · 2023
Earlier work this paper cites.
Universal and transferable adversarial attacks on aligned language models, 2023
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson · 2023
Earlier work this paper cites.
Jailbreaking leading safety-aligned llms with simple adaptive attacks, 2024
Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion · 2024
Earlier work this paper cites.
How (un)ethical are instruction-centric responses of llms? unveiling the vulnerabilities of safety guardrails to harmful queries, 2024
Somnath Banerjee, Sayan Layek, Rima Hazra, and Animesh Mukherjee · 2024
Earlier work this paper cites.
Play guessing game with llm: Indirect jailbreak attack with implicit clues, 2024
Zhiyuan Chang, Mingyang Li, Yi Liu, Junjie Wang, Qing Wang, and Yang Liu · 2024
Earlier work this paper cites.
Jailbreakbench: An open robustness benchmark for jailbreaking large language models, 2024
Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J. Pappas, Florian Tramer, Hamed Hassani, and Eric Wong · 2024
Earlier work this paper cites.
Red teaming gpt-4v: Are gpt-4v safe against uni/multi-modal jailbreak attacks?, 2024
Shuo Chen, Zhen Han, Bailan He, Zifeng Ding, Wenqian Yu, Philip Torr, Volker Tresp, and Jindong Gu · 2024
Earlier work this paper cites.
Comprehensive assessment of jailbreak attacks against llms, 2024
Junjie Chu, Yugeng Liu, Ziqing Yang, Xinyue Shen, Michael Backes, and Yang Zhang · 2024
Earlier work this paper cites.
Masterkey: Automated jailbreaking of large language model chatbots
Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu · 2024
Cited alongside, same era.
Pandora: Jailbreak gpts by retrieval augmented generation poisoning, 2024
Gelei Deng, Yi Liu, Kailong Wang, Yuekang Li, Tianwei Zhang, and Yang Liu · 2024
Cited alongside, same era.
Multilingual jailbreak challenges in large language models, 2024
Yue Deng, Wenxuan Zhang, Sinno Jialin Pan, and Lidong Bing · 2024
Cited alongside, same era.
A wolf in sheep’s clothing: Generalized nested jailbreak prompts can fool large language models easily, 2024
Peng Ding, Jun Kuang, Dan Ma, Xuezhi Cao, Yunsen Xian, Jiajun Chen, and Shujian Huang · 2024
Cited alongside, same era.
Analyzing the inherent response tendency of llms: Real-world instructions-driven jailbreak, 2024
Yanrui Du, Sendong Zhao, Ming Ma, Yuhan Chen, and Bing Qin · 2024
Cited alongside, same era.
Jailbreaking attack against multimodal large language model, 2024
Zhenxing Niu, Haodong Ren, Xinbo Gao, Gang Hua, and Rong Jin · 2024
Closest in time.
Visual adversarial examples jailbreak aligned large language models
Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Peter Henderson, Mengdi Wang, and Prateek Mittal · 2024
Closest in time.
Tricking LLMs into disobedience: Formalizing, analyzing, and detecting jailbreaks
Abhinav Sukumar Rao, Atharva Roshan Naik, Sachin Vashistha, Somak Aditya, and Monojit Choudhury · 2024
Closest in time.
Great, now write an article about that: The crescendo multi-turn llm jailbreak attack, 2024
Mark Russinovich, Ahmed Salem, and Ronen Eldan · 2024
Closest in time.
Xstest: A test suite for identifying exaggerated safety behaviours in large language models, 2024
Paul Röttger, Hannah Rose Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Coercing LLMs to do and reveal (almost) anything
Jonas Geiping, Alex Stein, Manli Shu, Khalid Saifullah, Yuxin Wen, and Tom Goldstein · 2024
Cited alongside, same era.
Aegis: Online adaptive ai content safety moderation with ensemble of llm experts, 2024
Shaona Ghosh, Prasoon Varshney, Erick Galinkin, and Christopher Parisien · 2024
Cited alongside, same era.
Cold-attack: Jailbreaking llms with stealthiness and controllability, 2024
Xingang Guo, Fangxu Yu, Huan Zhang, Lianhui Qin, and Bin Hu · 2024
Cited alongside, same era.
Pruning for protection: Increasing jailbreak resistance in aligned llms without fine-tuning, 2024
Adib Hasan, Ileana Rugina, and Alex Wang · 2024
Cited alongside, same era.
Query-based adversarial prompt generation, 2024
Jonathan Hayase, Ema Borevkovic, Nicholas Carlini, Florian Tramèr, and Milad Nasr · 2024
Cited alongside, same era.
Sowing the wind, reaping the whirlwind: The impact of editing language models, 2024
Rima Hazra, Sayan Layek, Somnath Banerjee, and Soujanya Poria · 2024
Cited alongside, same era.
Catastrophic jailbreak of open-source LLMs via exploiting generation
Yangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li, and Danqi Chen · 2024
Cited alongside, same era.
Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models
Erfan Shayegani, Yue Dong, and Nael Abu-Ghazaleh · 2024
Closest in time.
Attackeval: How to evaluate the effectiveness of jailbreak attacking on large language models, 2024
Dong Shu, Mingyu Jin, Suiyuan Zhu, Beichen Wang, Zihao Zhou, Chong Zhang, and Yongfeng Zhang · 2024
Closest in time.
Pal: Proxy-guided black-box attack on large language models, 2024
Chawin Sitawarin, Norman Mu, David Wagner, and Alexandre Araujo · 2024
Closest in time.
A strongreject for empty jailbreaks, 2024
Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, and Sam Toyer · 2024
Closest in time.
All in how you ask for it: Simple black-box method for jailbreak attacks
Kazuhiro Takemoto · 2024
Closest in time.
Meta llama guard 2
Llama Team · 2024
Closest in time.
From noise to clarity: Unraveling the adversarial suffix of large language model attacks via translation of text embeddings, 2024
Hao Wang, Hao Li, Minlie Huang, and Lei Sha · 2024
Closest in time.
Mitigating fine-tuning jailbreak attack with backdoor enhanced alignment, 2024
Jiongxiao Wang, Jiazhao Li, Yiquan Li, Xiangyu Qi, Junjie Hu, Yixuan Li, Patrick McDaniel, Muhao Chen, Bo Li, and Chaowei Xiao · 2024
Closest in time.
Adashield: Safeguarding multimodal large language models from structure-based attack via adaptive shield prompting, 2024
Yu Wang, Xiaogeng Liu, Yu Li, Muhao Chen, and Chaowei Xiao · 2024
Closest in time.
Jailbreaking gpt-4v via self-adversarial attacks with system prompts, 2024
Yuanwei Wu, Xiang Li, Yixin Liu, Pan Zhou, and Lichao Sun · 2024
Closest in time.
Tastle: Distract large language models for automatic jailbreak attack, 2024
Zeguan Xiao, Yan Yang, Guanhua Chen, and Yun Chen · 2024
Closest in time.
Safedecoding: Defending against jailbreak attacks via safety-aware decoding, 2024
Zhangchen Xu, Fengqing Jiang, Luyao Niu, Jinyuan Jia, Bill Yuchen Lin, and Radha Poovendran · 2024
Closest in time.
Low-resource languages jailbreak gpt-4, 2024
Zheng-Xin Yong, Cristina Menghini, and Stephen H. Bach · 2024
Closest in time.
Don’t listen to me: Understanding and exploring jailbreak prompts of large language models, 2024
Zhiyuan Yu, Xiaogeng Liu, Shunning Liang, Zach Cameron, Chaowei Xiao, and Ning Zhang · 2024
Closest in time.
GPT-4 is too smart to be safe: Stealthy chat with LLMs via cipher
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu · 2024
Closest in time.
Rigorllm: Resilient guardrails for large language models against undesired content, 2024
Zhuowen Yuan, Zidi Xiong, Yi Zeng, Ning Yu, Ruoxi Jia, Dawn Song, and Bo Li · 2024
Closest in time.
How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms, 2024
Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi · 2024
Closest in time.
Autodefense: Multi-agent llm defense against jailbreak attacks, 2024
Yifan Zeng, Yiran Wu, Xiao Zhang, Huazheng Wang, and Qingyun Wu · 2024
Closest in time.
Intention analysis makes llms a good jailbreak defender, 2024
Yuqi Zhang, Liang Ding, Lefei Zhang, and Dacheng Tao · 2024
Closest in time.
Psysafe: A comprehensive framework for psychological-based attack, defense, and evaluation of multi-agent system safety, 2024
Zaibin Zhang, Yongting Zhang, Lijun Li, Hongzhi Gao, Lijun Wang, Huchuan Lu, Feng Zhao, Yu Qiao, and Jing Shao · 2024
Closest in time.
Weak-to-strong jailbreaking on large language models, 2024
Xuandong Zhao, Xianjun Yang, Tianyu Pang, Chao Du, Lei Li, Yu-Xiang Wang, and William Yang Wang · 2024
Closest in time.
On prompt-driven safeguarding for large language models
Chujie Zheng, Fan Yin, Hao Zhou, Fandong Meng, Jie Zhou, Kai-Wei Chang, Minlie Huang, and Nanyun Peng · 2024
Closest in time.
Robust prompt optimization for defending language models against jailbreaking attacks, 2024
Andy Zhou, Bo Li, and Haohan Wang · 2024
Closest in time.
Easyjailbreak: A unified framework for jailbreaking large language models, 2024
Weikang Zhou, Xiao Wang, Limao Xiong, Han Xia, Yingshuang Gu, Mingxu Chai, Fukang Zhu, Caishuang Huang, Shihan Dou, Zhiheng Xi, Rui Zheng, Songyang Gao, Yicheng Zou, Hang Yan, Yifan Le, Ruohui Wang, Lijun Li, Jing Shao, Tao Gui, Qi Zhang, and Xuanjing Huang · 2024
Closest in time.
Don’t say no: Jailbreaking llm by suppressing refusal, 2024
Yukai Zhou and Wenjie Wang · 2024
Closest in time.
Safety fine-tuning at (almost) no cost: A baseline for vision large language models, 2024
Yongshuo Zong, Ondrej Bohdal, Tingyang Yu, Yongxin Yang, and Timothy Hospedales · 2024
Closest in time.
Is the system message really important to jailbreaks in large language models?, 2024
Xiaotian Zou, Yongkang Chen, and Ke Li · 2024
Closest in time.