Fetching the paper…
Reading the bibliography…
Recent large language model (LLM) defenses have greatly improved models' ability to refuse harmful queries, even when adversarially attacked.
Adversarial examples are not bugs, they are features, 2019
A. Ilyas, S. Santurkar, D. Tsipras, L. Engstrom, B. Tran, and A. Madry · 1905
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks, 2021
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W. tau Yih, T. Rocktäschel, S. Riedel, and D. Kiela · 2005
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
T. Shin, Y. Razeghi, R. L. Logan IV, E. Wallace, and S. Singh · 2010
Earlier work this paper cites.
Towards making systems forget with machine unlearning
Y. Cao and J. Yang · 2015
Earlier work this paper cites.
Explaining and harnessing adversarial examples, 2015
I. J. Goodfellow, J. Shlens, and C. Szegedy · 2015
Earlier work this paper cites.
Adversarial examples for evaluating reading comprehension systems
R. Jia and P. Liang · 2017
Earlier work this paper cites.
Adversarial examples in the physical world, 2017
A. Kurakin, I. Goodfellow, and S. Bengio · 2017
Earlier work this paper cites.
A. Athalye, N. Carlini, and D. Wagner · 2018
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks, 2019
A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu · 2019
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing NLP
E. Wallace, S. Feng, N. Kandpal, M. Gardner, and S. Singh · 2019
Earlier work this paper cites.
AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
T. Shin, Y. Razeghi, R. L. Logan IV, E. Wallace, and S. Singh · 2020
Earlier work this paper cites.
Machine unlearning
L. Bourtoule, V. Chandrasekaran, C. A. Choquette-Choo, H. Jia, A. Travers, B. Zhang, D. Lie, and N. Papernot · 2021
Earlier work this paper cites.
Unsolved problems in ml safety
D. Hendrycks, N. Carlini, J. Schulman, and J. Steinhardt · 2021
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
D. Ganguli, L. Lovitt, J. Kernion, A. Askell, Y. Bai, S. Kadavath, B. Mann, E. Perez, N. Schiefer, K. Ndousse, et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Earlier work this paper cites.
Red teaming language models with language models
E. Perez, S. Huang, F. Song, T. Cai, R. Ring, J. Aslanides, A. Glaese, N. McAleese, and G. Irving · 2022
Earlier work this paper cites.
Are aligned neural networks adversarially aligned?
N. Carlini, M. Nasr, C. A. Choquette-Choo, M. Jagielski, I. Gao, A. Awadalla, P. W. Koh, D. Ippolito, K. Lee, F. Tramer, et al · 2023
Earlier work this paper cites.
Explore, establish, exploit: Red teaming language models from scratch
S. Casper, J. Lin, J. Kwon, G. Culp, and D. Hadfield-Menell · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries
P. Chao, A. Robey, E. Dobriban, H. Hassani, G. J. Pappas, and E. Wong · 2023
Earlier work this paper cites.
P. Ding, J. Kuang, D. Ma, X. Cao, Y. Xian, J. Chen, and S. Huang · 2023
Earlier work this paper cites.
Mart: Improving llm safety with multi-round automatic red-teaming
S. Ge, C. Zhou, R. Hou, M. Khabsa, Y.-C. Wang, Q. Wang, J. Han, and Y. Mao · 2023
Earlier work this paper cites.
Llm censorship: A machine learning challenge or a computer security problem?
D. Glukhov, I. Shumailov, Y. Gal, N. Papernot, and V. Papyan · 2023
Earlier work this paper cites.
Red-teaming large language models to identify novel ai risks, 2023
W. House · 2023
Earlier work this paper cites.
Summon a demon and bind it: A grounded theory of llm red teaming in the wild, 2023
N. Inie, J. Stray, and L. Derczynski · 2023
Earlier work this paper cites.
Autodan: Generating stealthy jailbreak prompts on aligned large language models, 2023
X. Liu, N. Xu, M. Chen, and C. Xiao · 2023
Earlier work this paper cites.
Tree of attacks: Jailbreaking black-box llms automatically, 2023
A. Mehrotra, M. Zampetakis, P. Kassianik, B. Nelson, H. Anderson, Y. Singer, and A. Karbasi · 2023
Cited alongside, same era.
Gpt-4 technical report, 2023
OpenAI · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn · 2023
Cited alongside, same era.
Activation addition: Steering language models without optimization
A. Turner, L. Thiergart, D. Udell, G. Leech, U. Mini, and M. MacDiarmid · 2023
Cited alongside, same era.
Jailbroken: How does llm safety training fail?
A. Wei, N. Haghtalab, and J. Steinhardt · 2023
Cited alongside, same era.
Lora fine-tuning efficiently undoes safety training in llama 2-chat 70b, 2024
S. Lermen, C. Rogers-Smith, and J. Ladish · 2024
Closest in time.
Large language model unlearning via embedding-corrupted prompts
C. Y. Liu, Y. Wang, J. Flanigan, and Y. Liu · 2024
Closest in time.
Eight methods to evaluate robust unlearning in llms, 2024
A. Lynch, P. Guo, A. Ewart, S. Casper, and D. Hadfield-Menell · 2024
Closest in time.
Prp: Propagating universal perturbations to attack large language model guard-rails
N. Mangaokar, A. Hooda, J. Choi, S. Chandrashekaran, K. Fawaz, S. Jha, and A. Prakash · 2024
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Low-resource languages jailbreak gpt-4
Z.-X. Yong, C. Menghini, and S. H. Bach · 2023
Cited alongside, same era.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts, 2023
J. Yu, X. Lin, Z. Yu, and X. Xing · 2023
Cited alongside, same era.
Does refusal training in llms generalize to the past tense?, 2024
M. Andriushchenko and N. Flammarion · 2024
Cited alongside, same era.
Jailbreaking leading safety-aligned LLMs with simple adaptive attacks
M. Andriushchenko, F. Croce, and N. Flammarion · 2024
Cited alongside, same era.
Many-shot jailbreaking
C. Anil, E. Durmus, M. Sharma, J. Benton, S. Kundu, J. Batson, N. Rimsky, M. Tong, J. Mu, D. Ford, et al · 2024
Cited alongside, same era.
Unlearning via rmu is mostly shallow, 2024
A. Arditi and bilalchughtai · 2024
Cited alongside, same era.
Refusal in language models is mediated by a single direction, 2024
A. Arditi, O. Obeso, A. Syed, D. Paleka, N. Panickssery, W. Gurnee, and N. Nanda · 2024
Cited alongside, same era.
M. Mazeika, L. Phan, X. Yin, A. Zou, Z. Wang, N. Mu, E. Sakhaee, N. Li, S. Basart, B. Li, D. Forsyth, and D. Hendrycks · 2024
Closest in time.
Feedback loops with language models drive in-context reward hacking, 2024
A. Pan, E. Jones, M. Jagadeesan, and J. Steinhardt · 2024
Closest in time.
Steering llama 2 via contrastive activation addition, 2024
N. Panickssery, N. Gabrieli, J. Schulz, M. Tong, E. Hubinger, and A. M. Turner · 2024
Closest in time.
Safety alignment should be made more than just a few tokens deep, 2024
X. Qi, A. Panda, K. Lyu, X. Ma, S. Roy, A. Beirami, P. Mittal, and P. Henderson · 2024
Closest in time.
Representation noising effectively prevents harmful fine-tuning on llms
D. Rosati, J. Wehner, K. Williams, L. Bartoszcze, D. Atanasov, R. Gonzales, S. Majumdar, C. Maple, H. Sajjad, and F. Rudzicz · 2024
Closest in time.
Great, now write an article about that: The crescendo multi-turn llm jailbreak attack
M. Russinovich, A. Salem, and R. Eldan · 2024
Closest in time.
Revisiting the robust alignment of circuit breakers, 2024
L. Schwinn and S. Geisler · 2024
Closest in time.
"do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models, 2024
X. Shen, Z. Chen, M. Backes, Y. Shen, and Y. Zhang · 2024
Closest in time.
Targeted latent adversarial training improves robustness to persistent harmful behaviors in llms
A. Sheshadri, A. Ewart, P. Guo, A. Lynch, C. Wu, V. Hebbar, H. Sleight, A. C. Stickland, E. Perez, D. Hadfield-Menell, and S. Casper · 2024
Closest in time.
Pal: Proxy-guided black-box attack on large language models, 2024
C. Sitawarin, N. Mu, D. Wagner, and A. Araujo · 2024
Closest in time.
Multi-turn context jailbreak attack on large language models from first principles, 2024
X. Sun, D. Zhang, D. Yang, Q. Zou, and H. Li · 2024
Closest in time.
Tamper-resistant safeguards for open-weight llms, 2024
R. Tamirisa, B. Bharathi, L. Phan, A. Zhou, A. Gatti, T. Suresh, M. Lin, J. Wang, R. Wang, R. Arel, A. Zou, D. Song, B. Li, D. Hendrycks, and M. Mazeika · 2024
Closest in time.
Gemini: A family of highly capable multimodal models, 2024
G. Team et al · 2024
Closest in time.
Breaking circuit breakers, 2024
T. B. Thompson and M. Sklar · 2024
Closest in time.
Star: Sociotechnical approach to red teaming language models, 2024
L. Weidinger, J. Mellor, B. G. Pegueroles, N. Marchal, R. Kumar, K. Lum, C. Akbulut, M. Diaz, S. Bergman, M. Rodriguez, V. Rieser, and W. Isaac · 2024
Closest in time.
Efficient adversarial training in llms with continuous attacks, 2024
S. Xhonneux, A. Sordoni, S. Günnemann, G. Gidel, and L. Schwinn · 2024
Closest in time.
Refuse whenever you feel unsafe: Improving safety in llms via decoupled refusal training, 2024
Y. Yuan, W. Jiao, W. Wang, J. tse Huang, J. Xu, T. Liang, P. He, and Z. Tu · 2024
Closest in time.
Y. Zeng, H. Lin, J. Zhang, D. Yang, R. Jia, and W. Shi · 2024
Closest in time.
Robust prompt optimization for defending language models against jailbreaking attacks, 2024
A. Zhou, B. Li, and H. Wang · 2024
Closest in time.
Improving alignment and robustness with circuit breakers, 2024
A. Zou, L. Phan, J. Wang, D. Duenas, M. Lin, M. Andriushchenko, R. Wang, Z. Kolter, M. Fredrikson, and D. Hendrycks · 2024
Closest in time.