Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have demonstrated remarkable capabilities in natural language processing tasks, but their vulnerability to jailbreak attacks poses significant security risks.
Rlprompt: Optimizing discrete text prompts with reinforcement learning, 2022
Mingkai Deng, Jianyu Wang, Cheng-Ping Hsieh, Yihan Wang, Han Guo, Tianmin Shu, Meng Song, E. Xing, and Zhiting Hu · 2022
Earlier work this paper cites.
Cti4ai: Threat intelligence generation and sharing after red teaming ai models, 2022
C. Nguyen, Caleb Morgan, and Sudip Mittal · 2022
Earlier work this paper cites.
Red teaming language models with language models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and G. Irving · 2022
Earlier work this paper cites.
Ignore previous prompt: Attack techniques for language models, 2022
Fábio Perez and Ian Ribeiro · 2022
Earlier work this paper cites.
Promptattack: Prompt-based attack for language models via gradient search, 2022
Yundi Shi, Piji Li, Changchun Yin, Zhaoyang Han, Lu Zhou, and Zhe Liu · 2022
Earlier work this paper cites.
A prompting-based approach for adversarial example generation and robustness enhancement, 2022
Yuting Yang, Pei Huang, Juan Cao, Jintao Li, Yun Lin, J. Dong, Feifei Ma, and Jian Zhang · 2022
Earlier work this paper cites.
Red-teaming large language models using chain of utterances for safety-alignment, 2023
Rishabh Bhardwaj and Soujanya Poria · 2023
Earlier work this paper cites.
Defending against alignment-breaking attacks via robustly aligned llm, 2023
Bochuan Cao, Yu Cao, Lu Lin, and Jinghui Chen · 2023
Earlier work this paper cites.
Explore, establish, exploit: Red teaming language models from scratch, 2023
Stephen Casper, Jason Lin, Joe Kwon, Gatlen Culp, and Dylan Hadfield-Menell · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries, 2023
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong · 2023
Earlier work this paper cites.
She had cobalt blue eyes: Prompt testing to create aligned and sustainable language models, 2023
Veronica Chatrath, Oluwanifemi Bamgbose, and Shaina Raza · 2023
Earlier work this paper cites.
Jailbreaker in jail: Moving target defense for large language models, 2023
Bocheng Chen, Advait Paliwal, and Qiben Yan · 2023
Earlier work this paper cites.
A wolf in sheep’s clothing: Generalized nested jailbreak prompts can fool large language models easily, 2023
Peng Ding, Jun Kuang, Dan Ma, Xuezhi Cao, Yunsen Xian, Jiajun Chen, and Shujian Huang · 2023
Earlier work this paper cites.
Mart: Improving llm safety with multi-round automatic red-teaming, 2023
Suyu Ge, Chunting Zhou, Rui Hou, Madian Khabsa, Yi-Chia Wang, Qifan Wang, Jiawei Han, and Yuning Mao · 2023
Earlier work this paper cites.
Ai control: Improving safety despite intentional subversion, 2023
Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, and Fabien Roger · 2023
Earlier work this paper cites.
Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection, 2023
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, C. Endres, Thorsten Holz, and Mario Fritz · 2023
Earlier work this paper cites.
Codelmsec benchmark: Systematically evaluating and finding security vulnerabilities in black-box code language models, 2023
Hossein Hajipour, Thorsten Holz, Lea Schonherr, and Mario Fritz · 2023
Earlier work this paper cites.
Llm self defense: By self examination, llms know they are being tricked, 2023
Alec Helbling, Mansi Phute, Matthew Hull, and Duen Horng Chau · 2023
Earlier work this paper cites.
Token-level adversarial prompt detection based on perplexity measures and contextual information, 2023
Zhengmian Hu, Gang Wu, Saayan Mitra, Ruiyi Zhang, Tong Sun, Heng Huang, and Vishy Swaminathan · 2023
Earlier work this paper cites.
Catastrophic jailbreak of open-source llms via exploiting generation, 2023
Yangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li, and Danqi Chen · 2023
Earlier work this paper cites.
Baseline defenses for adversarial attacks against aligned language models, 2023
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein · 2023
Earlier work this paper cites.
Open sesame! universal black box jailbreaking of large language models, 2023
Raz Lapid, Ron Langberg, and Moshe Sipper · 2023
Earlier work this paper cites.
Query-efficient black-box red teaming via bayesian optimization, 2023
Deokjae Lee, JunYeong Lee, Jung-Woo Ha, Jin-Hwa Kim, Sang-Woo Lee, Hwaran Lee, and Hyun Oh Song · 2023
Earlier work this paper cites.
Evaluating the instruction-following robustness of large language models to prompt injection, 2023
Zekun Li, Baolin Peng, Pengcheng He, and Xifeng Yan · 2023
Earlier work this paper cites.
Student-teacher prompting for red teaming to improve guardrails, 2023
Rodrigo Revilla Llaca, Victoria Leskoschek, Vitor Costa Paiva, Cătălin Lupău, Philip Lippmann, and Jie Yang · 2023
Earlier work this paper cites.
Tree of attacks: Jailbreaking black-box llms automatically, 2023
Anay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum Anderson, Yaron Singer, and Amin Karbasi · 2023
Earlier work this paper cites.
Maatphor: Automated variant analysis for prompt injection attacks, 2023
Ahmed Salem, Andrew J. Paverd, and Boris Köpf · 2023
Earlier work this paper cites.
Ignore this title and hackaprompt: Exposing systemic vulnerabilities of llms through a global scale prompt hacking competition, 2023
Sander Schulhoff, Jeremy Pinto, Anaum Khan, Louis-Franccois Bouchard, Chenglei Si, Svetlina Anati, Valen Tagliabue, Anson Liu Kost, Christopher Carnahan, and Jordan L. Boyd-Graber · 2023
Earlier work this paper cites.
No offense taken: Eliciting offensiveness from language models, 2023
Anugya Srivastava, Rahul Ahuja, and Rohith Mukku · 2023
Earlier work this paper cites.
Cover: A heuristic greedy adversarial attack on prompt-based learning in language models, 2023
Zihao Tan, Qingliang Chen, Wenbin Zhu, and Yongjian Huang · 2023
Cited alongside, same era.
Evil geniuses: Delving into the safety of llm-based agents, 2023
Yu Tian, Xiao Yang, Jingyuan Zhang, Yinpeng Dong, and Hang Su · 2023
Cited alongside, same era.
Tensor trust: Interpretable prompt injection attacks from an online game, 2023
S. Toyer, Olivia Watkins, Ethan Mendes, Justin Svegliato, Luke Bailey, Tiffany Wang, Isaac Ong, Karim Elmaaroufi, Pieter Abbeel, Trevor Darrell, Alan Ritter, and Stuart Russell · 2023
Cited alongside, same era.
Self-deception: Reverse penetrating the semantic firewall of large language models, 2023
Zhenhua Wang, Wei Xie, Kai Chen, Baosheng Wang, Zhiwen Gui, and Enze Wang · 2023
Cited alongside, same era.
Jailbreaking gpt-4v via self-adversarial attacks with system prompts, 2023
Yuanwei Wu, Xiang Li, Yixin Liu, Pan Zhou, and Lichao Sun · 2023
Cited alongside, same era.
Can reinforcement learning unlock the hidden dangers in aligned large language models?, 2024
Mohammad Bahrami Karkevandi, Nishant Vishwamitra, and Peyman Najafirad · 2024
Closest in time.
Alpaca against vicuna: Using llms to uncover memorization of llms, 2024
Aly M. Kassem, Omar Mahmoud, Niloofar Mireshghallah, Hyunwoo Kim, Yulia Tsvetkov, Yejin Choi, Sherif Saad, and Santu Rana · 2024
Closest in time.
Prompt injection attacks in defended systems, 2024
Daniil Khomsky, Narek Maloyan, and Bulat Nutfullin · 2024
Closest in time.
Robust safety classifier against jailbreaking attacks: Adversarial prompt shield, 2024
Jinhwa Kim, Ali Derakhshan, and Ian G. Harris · 2024
Closest in time.
Towards detecting unanticipated bias in large language models, 2024
Anna Kruspe · 2024
Closest in time.
Fine-tuning, quantization, and llms: Navigating unintended outcomes, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fuzzllm: A novel and universal fuzzing framework for proactively discovering jailbreak vulnerabilities in large language models, 2023
Dongyu Yao, Jianshu Zhang, Ian G. Harris, and Marcel Carlsson · 2023
Cited alongside, same era.
Benchmarking and defending against indirect prompt injection attacks on large language models, 2023
Jingwei Yi, Yueqi Xie, Bin Zhu, Keegan Hines, Emre Kiciman, Guangzhong Sun, Xing Xie, and Fangzhao Wu · 2023
Cited alongside, same era.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts, 2023
Jiahao Yu, Xingwei Lin, Zheng Yu, and Xinyu Xing · 2023
Cited alongside, same era.
Are you still on track!? catching llm task drift with activations, 2024
Sahar Abdelnabi, Aideen Fay, Giovanni Cherubin, Ahmed Salem, Mario Fritz, and Andrew Paverd · 2024
Cited alongside, same era.
Prompt leakage effect and defense strategies for multi-turn llm interactions, 2024
Divyansh Agarwal, A. R. Fabbri, Philippe Laban, Shafiq R. Joty, Caiming Xiong, and Chien-Sheng Wu · 2024
Cited alongside, same era.
Automatic pseudo-harmful prompt generation for evaluating false refusals in large language models, 2024
Bang An, Sicheng Zhu, Ruiyi Zhang, Michael-Andrei Panaitescu-Liess, Yuancheng Xu, and Furong Huang · 2024
Cited alongside, same era.
Jailbreaking leading safety-aligned llms with simple adaptive attacks, 2024
Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion · 2024
Cited alongside, same era.
Divyanshu Kumar, Anurakt Kumar, Sahil Agarwal, and P. Harshangi · 2024
Closest in time.
Learning diverse attacks on large language models for robust red-teaming and safety tuning, 2024
Seanie Lee, Minsu Kim, Lynn Cherif, David Dobre, Juho Lee, Sung Ju Hwang, Kenji Kawaguchi, Gauthier Gidel, Y. Bengio, Nikolay Malkin, and Moksh Jain · 2024
Closest in time.
Why are my prompts leaked? unraveling prompt extraction threats in customized large language models, 2024
Zi Liang, Haibo Hu, Qingqing Ye, Yaxin Xiao, and Haoyang Li · 2024
Closest in time.
Autojailbreak: Exploring jailbreak attacks and defenses through a dependency lens, 2024
Lin Lu, Hai Yan, Zenghui Yuan, Jiawen Shi, Wenqi Wei, Pin-Yu Chen, and Pan Zhou · 2024
Closest in time.
Adappa: Adaptive position pre-fill jailbreak attack approach targeting llms, 2024
Lijia Lv, Weigang Zhang, Xuehai Tang, Jie Wen, Feng Liu, Jizhong Han, and Songlin Hu · 2024
Closest in time.
Prp: Propagating universal perturbations to attack large language model guard-rails, 2024
Neal Mangaokar, Ashish Hooda, Jihye Choi, Shreyas Chandrashekaran, Kassem Fawaz, Somesh Jha, and Atul Prakash · 2024
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal, 2024
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks · 2024
Closest in time.
Kov: Transferable and naturalistic black-box llm attacks using markov decision processes and tree search, 2024
Robert J. Moss · 2024
Closest in time.
Efficient llm-jailbreaking by introducing visual modality, 2024
Zhenxing Niu, Yuyao Sun, Haodong Ren, Haoxuan Ji, Quan Wang, Xiaoke Ma, Gang Hua, and Rong Jin · 2024
Closest in time.
Ferret: Faster and effective automated red teaming with reward-based scoring technique, 2024
Tej Deep Pala, Vernon Y.H. Toh, Rishabh Bhardwaj, and Soujanya Poria · 2024
Closest in time.
Advprompter: Fast adaptive adversarial prompting for llms, 2024
Anselm Paulus, Arman Zharmagambetov, Chuan Guo, Brandon Amos, and Yuandong Tian · 2024
Closest in time.
Learning to poison large language models during instruction tuning, 2024
Yao Qiang, Xiangyu Zhou, Saleh Zare Zade, Mohammad Amin Roshani, Douglas Zytko, and Dongxiao Zhu · 2024
Closest in time.
Gpt-4 jailbreaks itself with near-perfect success using self-explanation, 2024
Govind Ramesh, Yao Dou, and Wei Xu · 2024
Closest in time.
Rainbow teaming: Open-ended generation of diverse adversarial prompts, 2024
Mikayel Samvelyan, S. Raparthy, Andrei Lupu, Eric Hambro, Aram H. Markosyan, Manish Bhatt, Yuning Mao, Minqi Jiang, Jack Parker-Holder, Jakob Foerster, Tim Rocktaschel, and Roberta Raileanu · 2024
Closest in time.
Rapid optimization for jailbreaking llms via subconscious exploitation and echopraxia, 2024
Guangyu Shen, Siyuan Cheng, Kai xian Zhang, Guanhong Tao, Shengwei An, Lu Yan, Zhuo Zhang, Shiqing Ma, and Xiangyu Zhang · 2024
Closest in time.
Latent adversarial training improves robustness to persistent harmful behaviors in llms, 2024
Abhay Sheshadri, Aidan Ewart, Phillip Guo, Aengus Lynch, Cindy Wu, Vivek Hebbar, Henry Sleight, Asa Cooper Stickland, Ethan Perez, Dylan Hadfield-Menell, and Stephen Casper · 2024
Closest in time.
Pal: Proxy-guided black-box attack on large language models, 2024
Chawin Sitawarin, Norman Mu, David Wagner, and Alexandre Araujo · 2024
Closest in time.
Multi-turn context jailbreak attack on large language models from first principles, 2024
Xiongtao Sun, Deyue Zhang, Dongdong Yang, Quanchen Zou, and Hui Li · 2024
Closest in time.
Prompt evolution through examples for large language models–a case study in game comment toxicity classification, 2024
Pittawat Taveekitworachai, Febri Abdullah, Mustafa Can Gursesli, Antonio Lanatà, Andrea Guazzini, and R. Thawonmas · 2024
Closest in time.
Gradient-based language model red teaming, 2024
Nevan Wichers, Carson E. Denison, and Ahmad Beirami · 2024
Closest in time.
Tastle: Distract large language models for automatic jailbreak attack, 2024
Zeguan Xiao, Yan Yang, Guanhua Chen, and Yun Chen · 2024
Closest in time.
Llm-fuzzer: Scaling assessment of large language model jailbreaks, 2024
Jiahao Yu, Xingwei Lin, Zheng Yu, and Xinyu Xing · 2024
Closest in time.
Round trip translation defence against large language model jailbreaking attacks, 2024
Canaan Yung, H. M. Dolatabadi, S. Erfani, and Christopher Leckie · 2024
Closest in time.
Dpp-based adversarial prompt searching for lanugage models, 2024
Xu Zhang and Xiaojun Wan · 2024
Closest in time.