Fetching the paper…
Reading the bibliography…
This paper studies the vulnerabilities of transformer-based Large Language Models (LLMs) to jailbreaking attacks, focusing specifically on the optimization-based Greedy Coordinate Gradient (GCG) strategy.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, and et al · 2020
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, and et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, and et al · 2022
Earlier work this paper cites.
Jailbreak chat
Alex Albert · 2023
Earlier work this paper cites.
Are aligned neural networks adversarially aligned?
Nicholas Carlini, Milad Nasr, Christopher A Choquette-Choo, Matthew Jagielski, Irena Gao, Pang Wei W Koh, Daphne Ippolito, Florian Tramer, and Ludwig Schmidt · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong · 2023
Earlier work this paper cites.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, and et al · 2023
Earlier work this paper cites.
Open sesame! universal black box jailbreaking of large language models
Raz Lapid, Ron Langberg, and Moshe Sipper · 2023
Earlier work this paper cites.
Scalable and transferable black-box jailbreaks for language models via persona modulation
Rusheb Shah, Quentin Feuillade-Montixi, Soroush Pour, Arush Tagade, Stephen Casper, and Javier Rando · 2023
Earlier work this paper cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, and et al · 2023
Earlier work this paper cites.
Sight beyond text: Multi-modal training enhances llms in truthfulness and ethics
Haoqin Tu, Bingchen Zhao, Chen Wei, and Cihang Xie · 2023
Earlier work this paper cites.
Jailbroken: How does LLM safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Cited alongside, same era.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts
Jiahao Yu, Xingwei Lin, Zheng Yu, and Xinyu Xing · 2023
Cited alongside, same era.
Autodan: Interpretable gradient-based adversarial attacks on large language models
Sicheng Zhu, Ruiyi Zhang, Bang An, Gang Wu, Joe Barrow, Zichao Wang, Furong Huang, Ani Nenkova, and Tong Sun · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson · 2023
Cited alongside, same era.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, and et al · 2024
Closest in time.
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, and et al · 2024
Closest in time.
Zeyi Liao and Huan Sun · 2024
Closest in time.
Autodan: Generating stealthy jailbreak prompts on aligned large language models
Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao · 2024
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, and et al · 2024
Cited alongside, same era.
Jailbreaking leading safety-aligned llms with simple adaptive attacks
Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion · 2024
Cited alongside, same era.
Introducing the next generation of claude
Anthropic · 2024
Cited alongside, same era.
The hacking of chatgpt is just getting started
Matt Burgess · 2024
Cited alongside, same era.
Amazing “jailbreak” bypasses chatgpt’s ethics safeguards
Jon Christian · 2024
Cited alongside, same era.
Masterkey: Automated jailbreaking of large language model chatbots
Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu · 2024
Cited alongside, same era.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, and et al · 2024
Cited alongside, same era.
Attacking large language models with projected gradient descent
Simon Geisler, Tom Wollschläger, M. H. I. Abdalla, Johannes Gasteiger, and Stephan Günnemann · 2024
Cited alongside, same era.
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks · 2024
Closest in time.
Tree of attacks: Jailbreaking black-box llms automatically
Anay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum Anderson, Yaron Singer, and Amin Karbasi · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, and et al · 2024
Closest in time.
All in how you ask for it: Simple black-box method for jailbreak attacks
Kazuhiro Takemoto · 2024
Closest in time.
Dan is my new friend
walkerspider · 2024
Closest in time.
Jailbreak and guard aligned language models with only few in-context demonstrations
Zeming Wei, Yifei Wang, and Yisen Wang · 2024
Closest in time.
Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu · 2024
Closest in time.
Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi · 2024
Closest in time.