Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) can be used to red team other models (e.g.
Intriguing properties of neural networks
C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus · 2014
Earlier work this paper cites.
Towards making systems forget with machine unlearning
Y. Cao and J. Yang · 2015
Earlier work this paper cites.
Explaining and harnessing adversarial examples, 2015
I. J. Goodfellow, J. Shlens, and C. Szegedy · 2015
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu · 2017
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
T. Shin, Y. Razeghi, R. L. Logan IV, E. Wallace, and S. Singh · 2020
Earlier work this paper cites.
Machine unlearning
L. Bourtoule, V. Chandrasekaran, C. A. Choquette-Choo, H. Jia, A. Travers, B. Zhang, D. Lie, and N. Papernot · 2021
Earlier work this paper cites.
Explore, establish, exploit: Red teaming language models from scratch
S. Casper, J. Lin, J. Kwon, G. Culp, and D. Hadfield-Menell · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries
P. Chao, A. Robey, E. Dobriban, H. Hassani, G. J. Pappas, and E. Wong · 2023
Earlier work this paper cites.
P. Ding, J. Kuang, D. Ma, X. Cao, Y. Xian, J. Chen, and S. Huang · 2023
Earlier work this paper cites.
Mart: Improving llm safety with multi-round automatic red-teaming
S. Ge, C. Zhou, R. Hou, M. Khabsa, Y.-C. Wang, Q. Wang, J. Han, and Y. Mao · 2023
Earlier work this paper cites.
Autodan: Generating stealthy jailbreak prompts on aligned large language models, 2023
X. Liu, N. Xu, M. Chen, and C. Xiao · 2023
Earlier work this paper cites.
Black box adversarial prompting for foundation models, 2023
N. Maus, P. Chao, E. Wong, and J. Gardner · 2023
Earlier work this paper cites.
Tree of attacks: Jailbreaking black-box llms automatically, 2023
A. Mehrotra, M. Zampetakis, P. Kassianik, B. Nelson, H. Anderson, Y. Singer, and A. Karbasi · 2023
Earlier work this paper cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Earlier work this paper cites.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts, 2023
J. Yu, X. Lin, Z. Yu, and X. Xing · 2023
Earlier work this paper cites.
Jailbreaking leading safety-aligned LLMs with simple adaptive attacks
M. Andriushchenko, F. Croce, and N. Flammarion · 2024
Cited alongside, same era.
Many-shot jailbreaking
C. Anil, E. Durmus, M. Sharma, J. Benton, S. Kundu, J. Batson, N. Rimsky, M. Tong, J. Mu, D. Ford, et al · 2024
Cited alongside, same era.
Introducing claude 3.5 sonnet, 2024
Anthropic · 2024
Cited alongside, same era.
Unlearning via rmu is mostly shallow, 2024
A. Arditi and bilalchughtai · 2024
Cited alongside, same era.
Refusal in language models is mediated by a single direction, 2024
A. Arditi, O. Obeso, A. Syed, D. Paleka, N. Panickssery, W. Gurnee, and N. Nanda · 2024
Cited alongside, same era.
Representation noising effectively prevents harmful fine-tuning on llms
D. Rosati, J. Wehner, K. Williams, L. Bartoszcze, D. Atanasov, R. Gonzales, S. Majumdar, C. Maple, H. Sajjad, and F. Rudzicz · 2024
Later among the works it cites.
Great, now write an article about that: The crescendo multi-turn llm jailbreak attack
M. Russinovich, A. Salem, and R. Eldan · 2024
Later among the works it cites.
Rainbow teaming: Open-ended generation of diverse adversarial prompts
M. Samvelyan, S. C. Raparthy, A. Lupu, E. Hambro, A. H. Markosyan, M. Bhatt, Y. Mao, M. Jiang, J. Parker-Holder, J. N. Foerster, T. Rocktäschel, and R. Raileanu · 2024
Later among the works it cites.
Revisiting the robust alignment of circuit breakers, 2024
L. Schwinn and S. Geisler · 2024
Later among the works it cites.
Targeted latent adversarial training improves robustness to persistent harmful behaviors in llms
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Beutel, K. Xiao, J. Heidecke, and L. Weng · 2024
Cited alongside, same era.
Gemini: A family of highly capable multimodal models, 2024
Google · 2024
Cited alongside, same era.
Endless jailbreaks with bijection learning, 2024
B. R. Y. Huang, M. Li, and L. Tang · 2024
Cited alongside, same era.
J. Hughes, S. Price, A. Lynch, R. Schaeffer, F. Barez, S. Koyejo, H. Sleight, E. Jones, E. Perez, and M. Sharma · 2024
Cited alongside, same era.
Babilong: Testing the limits of llms with long context reasoning-in-a-haystack
Y. Kuratov, A. Bulatov, P. Anokhin, I. Rodkin, D. Sorokin, A. Sorokin, and M. Burtsev · 2024
Cited alongside, same era.
Lora fine-tuning efficiently undoes safety training in llama 2-chat 70b, 2024
S. Lermen, C. Rogers-Smith, and J. Ladish · 2024
Cited alongside, same era.
Large language model unlearning via embedding-corrupted prompts
C. Y. Liu, Y. Wang, J. Flanigan, and Y. Liu · 2024
Cited alongside, same era.
A. Sheshadri, A. Ewart, P. Guo, A. Lynch, C. Wu, V. Hebbar, H. Sleight, A. C. Stickland, E. Perez, D. Hadfield-Menell, and S. Casper · 2024
Later among the works it cites.
Multi-turn context jailbreak attack on large language models from first principles, 2024
X. Sun, D. Zhang, D. Yang, Q. Zou, and H. Li · 2024
Later among the works it cites.
Tamper-resistant safeguards for open-weight llms, 2024
R. Tamirisa, B. Bharathi, L. Phan, A. Zhou, A. Gatti, T. Suresh, M. Lin, J. Wang, R. Wang, R. Arel, A. Zou, D. Song, B. Li, D. Hendrycks, and M. Mazeika · 2024
Later among the works it cites.
Efficient adversarial training in llms with continuous attacks, 2024
S. Xhonneux, A. Sordoni, S. Günnemann, G. Gidel, and L. Schwinn · 2024
Later among the works it cites.
Agentless: Demystifying llm-based software engineering agents, 2024
C. S. Xia, Y. Deng, S. Dunn, and L. Zhang · 2024
Later among the works it cites.
How johnny can persuade LLMs to jailbreak them: Rethinking persuasion to challenge AI safety by humanizing LLMs
Y. Zeng, H. Lin, J. Zhang, D. Yang, R. Jia, and W. Shi · 2024
Later among the works it cites.
Robust prompt optimization for defending language models against jailbreaking attacks, 2024
A. Zhou, B. Li, and H. Wang · 2024
Later among the works it cites.
Improving alignment and robustness with circuit breakers, 2024
A. Zou, L. Phan, J. Wang, D. Duenas, M. Lin, M. Andriushchenko, R. Wang, Z. Kolter, M. Fredrikson, and D. Hendrycks · 2024
Later among the works it cites.
Gemini 2.5: Our most intelligent ai model, Mar 2025
G. Deepmind · 2025
Closest in time.
Adversarial reasoning at jailbreaking time, 2025
M. Sabbaghi, P. Kassianik, G. Pappas, Y. Singer, A. Karbasi, and H. Hassani · 2025
Closest in time.
AIR-BENCH 2024: A safety benchmark based on regulation and policies specified risk categories
Y. Zeng, Y. Yang, A. Zhou, J. Z. Tan, Y. Tu, Y. Mai, K. Klyman, M. Pan, R. Jia, D. Song, P. Liang, and B. Li · 2025
Closest in time.