Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have achieved remarkable success across diverse tasks, yet they remain vulnerable to adversarial attacks, notably the well-known jailbreak attack.
Intriguing properties of neural networks
C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples
I. J. Goodfellow, J. Shlens, and C. Szegedy · 2014
Earlier work this paper cites.
Adversarial examples are not easily detected: Bypassing ten detection methods, 2017
N. Carlini and D. Wagner · 2017
Earlier work this paper cites.
Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples
A. Athalye, N. Carlini, and D. Wagner · 2018
Earlier work this paper cites.
Boosting adversarial attacks with momentum, 2018
Y. Dong et al · 2018
Earlier work this paper cites.
Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks
F. Croce and M. Hein · 2020
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts, 2020
T. Shin et al · 2020
Earlier work this paper cites.
Gradient-based adversarial attacks against text transformers, 2021
C. Guo et al · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback, 2022
Y. Bai et al · 2022
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022
Y. Bai et al · 2022
Earlier work this paper cites.
On the opportunities and risks of foundation models, 2022
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, and et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback, 2022
L. Ouyang et al · 2022
Earlier work this paper cites.
Combating misinformation in the age of llms: Opportunities and challenges
C. Chen and K. Shu · 2023
Earlier work this paper cites.
Robust classification via a single diffusion model
H. Chen, Y. Dong, Z. Wang, X. Yang, C. Duan, H. Su, and J. Zhu · 2023
Earlier work this paper cites.
Rethinking model ensemble in transfer-based adversarial attacks
H. Chen, Y. Zhang, Y. Dong, and J. Zhu · 2023
Earlier work this paper cites.
How robust is google’s bard to adversarial image attacks?
Y. Dong, H. Chen, J. Chen, Z. Fang, X. Yang, Y. Zhang, Y. Tian, H. Su, and J. Zhu · 2023
Cited alongside, same era.
Baseline defenses for adversarial attacks against aligned language models, 2023
N. Jain et al · 2023
Cited alongside, same era.
Mistral 7b, 2023
A. Q. Jiang et al · 2023
Cited alongside, same era.
Deepinception: Hypnotize large language model to be jailbreaker
X. Li et al · 2023
Cited alongside, same era.
Rain: Your language models can align themselves without finetuning, 2023
Y. Li et al · 2023
Cited alongside, same era.
Towards trustworthy and aligned machine learning: A data-centric survey with causality perspectives
Cognitive overload: Jailbreaking large language models with overloaded logical thinking
N. Xu et al · 2023
Later among the works it cites.
A survey on large language model (llm) security and privacy: The good, the bad, and the ugly
Y. Yao et al · 2023
Later among the works it cites.
Low-resource languages jailbreak gpt-4, 2023
Z.-X. Yong, C. Menghini, and S. H. Bach · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica · 2023
Later among the works it cites.
Autodan: Automatic and interpretable adversarial attacks on large language models
S. Zhu et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Liu, M. Chaudhary, and H. Wang · 2023
Cited alongside, same era.
Autodan: Generating stealthy jailbreak prompts on aligned large language models, 2023
X. Liu, N. Xu, M. Chen, and C. Xiao · 2023
Cited alongside, same era.
Jailbreaking chatgpt via prompt engineering: An empirical study, 2023
Y. Liu, G. Deng, Z. Xu, Y. Li, Y. Zheng, Y. Zhang, L. Zhao, T. Zhang, and Y. Liu · 2023
Cited alongside, same era.
Fine-tuning aligned language models compromises safety, even when users do not intend to!, 2023
X. Qi, Y. Zeng, T. Xie, P.-Y. Chen, R. Jia, P. Mittal, and P. Henderson · 2023
Cited alongside, same era.
M. A. Shah, R. Sharma, H. Dhamyal, R. Olivier, A. Shah, J. Konan, D. Alharthi, H. T. Bukhari, M. Baali, S. Deshmukh, et al · 2023
Cited alongside, same era.
"do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models, 2023
X. Shen, Z. Chen, M. Backes, Y. Shen, and Y. Zhang · 2023
Cited alongside, same era.
Jailbreak and guard aligned language models with only few in-context demonstrations
Z. Wei, Y. Wang, and Y. Wang · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models, 2023
A. Zou et al · 2023
Later among the works it cites.
Your diffusion model is secretly a certifiably robust classifier
H. Chen, Y. Dong, S. Shao, Z. Hao, X. Yang, H. Su, and J. Zhu · 2024
Closest in time.
Improved techniques for optimization-based jailbreaking on large language models
X. Jia et al · 2024
Closest in time.
Exploiting the index gradients for optimization-based jailbreaking on large language models
J. Li, Y. Hao, H. Xu, X. Wang, and Y. Hong · 2024
Closest in time.
How language models defend themselves against jailbreak: Theoretical understanding of self-correction through in-context learning
Y. Wang, Y. Wu, Z. Wei, S. Jegelka, and Y. Wang · 2024
Closest in time.
GPT-4 is too smart to be safe: Stealthy chat with LLMs via cipher
Y. Yuan, W. Jiao, W. Wang, J. tse Huang, P. He, S. Shi, and Z. Tu · 2024
Closest in time.
How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms, 2024
Y. Zeng et al · 2024
Closest in time.
Towards general conceptual model editing via adversarial representation engineering
Y. Zhang, Z. Wei, J. Sun, and M. Sun · 2024
Closest in time.
Towards the worst-case robustness of large language models
H. Chen, Y. Dong, Z. Wei, H. Su, and J. Zhu · 2025
Closest in time.