Fetching the paper…
Reading the bibliography…
Jailbreaks have been a central focus of research regarding the safety and reliability of large language models (LLMs), yet the mechanisms underlying these attacks remain poorly understood.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson. 2023 · 2023
Earlier work this paper cites.
Rethinking jailbreaking through the lens of representation engineering
Tianlong Li, Shihan Dou, Wenhao Liu, Muling Wu, Changze Lv, Rui Zheng, Xiaoqing Zheng, and Xuanjing Huang. 2024 · 2024
Cited alongside, same era.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio
Cited in the paper.
Many-shot jailbreaking
Cem Anil, Esin Durmus, Nina Panickssery, Mrinank Sharma, Joe Benton, Sandipan Kundu, Joshua Batson, Meg Tong, Jesse Mu, Daniel Ford, Fracesco Mosconi, Rajashree Agrawal, Rylan Schaeffer, Naomi Bashkansky, Samuel Svenningsen, Mike Lambert, Ansh Radhakrishnan, Carson Denison, Evan J. Hubinger, and 15 others
Cited in the paper.
Foundational challenges in assuring alignment and safety of large language models
Usman Anwar, Abulhair Saparov, Javier Rando, Daniel Paleka, Miles Turpin, Peter Hase, Ekdeep Singh Lubana, Erik Jenner, Stephen Casper, Oliver Sourbut, Benjamin L. Edelman, Zhaowei Zhang, Mario Günther, Anton Korinek, Jose Hernandez-Orallo, Lewis Hammond, Eric Bigelow, Alexander Pan, Lauro Langosco, and 23 others
Cited in the paper.
Refusal in language models is mediated by a single direction
Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, and Neel Nanda
Cited in the paper.
Probing pre-trained language models for cross-cultural differences in values
Arnav Arora, Lucie-Aimée Kaffee, and Isabelle Augenstein
Cited in the paper.
Understanding jailbreak success: A study of latent space dynamics in large language models
Sarah Ball, Frauke Kreuter, and Nina Rimsky
Cited in the paper.
Probing classifiers: Promises, shortcomings, and advances
Yonatan Belinkov
Cited in the paper.
Towards inference-time category-wise safety steering for large language models
Amrita Bhattacharjee, Shaona Ghosh, Traian Rebedea, and Christopher Parisien
Cited in the paper.
Understanding (un)reliability of steering vectors in language models
Joschka Braun, Carsten Eickhoff, David Krueger, Seyed Ali Bahrainian, and Dmitrii Krasheninnikov
Cited in the paper.
Discovering latent knowledge in language models without supervision
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt
Cited in the paper.
Breaking down the defenses: A comparative survey of attacks on large language models
Arijit Ghosh Chowdhury, Md Mofijul Islam, Vaibhav Kumar, Faysal Hossain Shezan, Vaibhav Kumar, Vinija Jain, and Aman Chadha
Cited in the paper.
Representation engineering: A top-down approach to ai transparency
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, and 2 others. 2025 · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…