Fetching the paper…
Reading the bibliography…
Most jailbreak papers claim the jailbreaks they propose are highly effective, often boasting near-100% attack success rates.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Y. Chen, H. Gao, G. Cui, F. Qi, L. Huang, Z. Liu, and M. Sun · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
D. Ganguli, L. Lovitt, J. Kernion, A. Askell, Y. Bai, S. Kadavath, B. Mann, E. Perez, N. Schiefer, K. Ndousse, et al · 2022
Earlier work this paper cites.
Red teaming language models with language models
E. Perez, S. Huang, F. Song, T. Cai, R. Ring, J. Aslanides, A. Glaese, N. McAleese, and G. Irving · 2022
Earlier work this paper cites.
On second thought, let’s not think step by step! bias and toxicity in zero-shot reasoning
O. Shaikh, H. Zhang, W. Held, M. Bernstein, and D. Yang · 2022
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Earlier work this paper cites.
Shield and spear: Jailbreaking aligned LLMs with generative prompting
Anonymous authors · 2023
Earlier work this paper cites.
Image hijacks: Adversarial images can control generative models at runtime
L. Bailey, E. Ong, S. Russell, and S. Emmons · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries
P. Chao, A. Robey, E. Dobriban, H. Hassani, G. J. Pappas, and E. Wong · 2023
Earlier work this paper cites.
Regulation of the european parliament and of the council laying down harmonised rules on artificial intelligence (artificial intelligence act) and amending certain union legislative acts, amendment 102
C. o. t. E. U. European Parliament · 2023
Earlier work this paper cites.
PICT: A zero-shot prompt template to automate evaluation, 2023
Q. Feuillade-Montixi · 2023
Earlier work this paper cites.
Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection
K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz · 2023
Earlier work this paper cites.
Catastrophic jailbreak of open-source LLMs via exploiting generation
Y. Huang, S. Gupta, M. Xia, K. Li, and D. Chen · 2023
Earlier work this paper cites.
Exploiting programmatic behavior of LLMs: Dual-use through standard security attacks
D. Kang, X. Li, I. Stoica, C. Guestrin, M. Zaharia, and T. Hashimoto · 2023
Cited alongside, same era.
Open sesame! universal black box jailbreaking of large language models
R. Lapid, R. Langberg, and M. Sipper · 2023
Cited alongside, same era.
Tree of attacks: Jailbreaking black-box LLMs automatically
A. Mehrotra, M. Zampetakis, P. Kassianik, B. Nelson, H. Anderson, Y. Singer, and A. Karbasi · 2023
Cited alongside, same era.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
X. Qi, Y. Zeng, T. Xie, P.-Y. Chen, R. Jia, P. Mittal, and P. Henderson · 2023
Cited alongside, same era.
SmoothLLM: Defending large language models against jailbreaking attacks
Low-resource languages jailbreak GPT-4
Z.-X. Yong, C. Menghini, and S. H. Bach · 2023
Later among the works it cites.
GPTFuzzer: Red teaming large language models with auto-generated jailbreak prompts
J. Yu, X. Lin, and X. Xing · 2023
Later among the works it cites.
Removing RLHF protections in GPT-4 via fine-tuning
Q. Zhan, R. Fang, R. Bindu, A. Gupta, T. Hashimoto, and D. Kang · 2023
Later among the works it cites.
AutoDAN: Automatic and interpretable adversarial attacks on large language models
S. Zhu, R. Zhang, B. An, G. Wu, J. Barrow, Z. Wang, F. Huang, A. Nenkova, and T. Sun · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Robey, E. Wong, H. Hassani, and G. J. Pappas · 2023
Cited alongside, same era.
X. Shen, Z. Chen, M. Backes, Y. Shen, and Y. Zhang · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models, 2023
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M.-A. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom · 2023
Cited alongside, same era.
Jailbroken: How does LLM safety training fail?
A. Wei, N. Haghtalab, and J. Steinhardt · 2023
Cited alongside, same era.
Voluntary AI commitments
White House · 2023
Cited alongside, same era.
Fact sheet: Biden-harris administration secures voluntary commitments from leading artificial intelligence companies to manage the risks posed by AI
White House Briefing Room · 2023
Cited alongside, same era.
You can use GPT-4 to create prompt injections against GPT-4, 2023
WitchBot · 2023
Cited alongside, same era.
Cognitive overload: Jailbreaking large language models with overloaded logical thinking
N. Xu, F. Wang, B. Zhou, B. Z. Li, C. Xiao, and M. Chen · 2023
Cited alongside, same era.
A. Zou, Z. Wang, J. Z. Kolter, and M. Fredrikson · 2023
Later among the works it cites.
dolphin-2.6-mixtral-8x7b
CognitiveComputations · 2024
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal, 2024
M. Mazeika, L. Phan, X. Yin, A. Zou, Z. Wang, N. Mu, E. Sakhaee, N. Li, S. Basart, B. Li, D. Forsyth, and D. Hendrycks · 2024
Closest in time.
Llama-3.1 system card
Meta · 2024
Closest in time.
GPT-3 API [text-davinci-003]
OpenAI · 2024
Closest in time.
Gpt-4o system card
OpenAI · 2024
Closest in time.
Generating terror: The risks of generative AI exploitation
G. Weimann, A. T. Pack, R. Sulciner, J. Scheinin, G. Rapaport, and D. Diaz · 2024
Closest in time.
Y. Zeng, H. Lin, J. Zhang, D. Yang, R. Jia, and W. Shi · 2024
Closest in time.