Fetching the paper…
Reading the bibliography…
In recent years, Large Language Models (LLMs) have gained widespread use, raising concerns about their security.
Actor-critic algorithms
V. Konda and J. Tsitsiklis · 1999
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
L. Li, W. Chu, J. Langford, and R. E. Schapire · 2010
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih · 2013
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
R. Lowe, Y. Wu, A. Tamar, J. Harb, P. Abbeel, and I. Mordatch · 2017
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al · 2022
Earlier work this paper cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, et al · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries, 2023
P. Chao, A. Robey, E. Dobriban, H. Hassani, G. J. Pappas, and E. Wong · 2023
Earlier work this paper cites.
Tree of attacks: Jailbreaking black-box llms automatically
A. Mehrotra, M. Zampetakis, P. Kassianik, B. Nelson, H. Anderson, Y. Singer, and A. Karbasi · 2023
Earlier work this paper cites.
Jailbroken: How does LLM safety training fail?
A. Wei, N. Haghtalab, and J. Steinhardt · 2023
Earlier work this paper cites.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts
J. Yu, X. Lin, and X. Xing · 2023
Cited alongside, same era.
Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher
Y. Yuan, W. Jiao, W. Wang, J.-t. Huang, P. He, S. Shi, and Z. Tu · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
A. Zou, Z. Wang, N. Carlini, M. Nasr, J. Z. Kolter, and M. Fredrikson · 2023
Cited alongside, same era.
Comprehensive assessment of jailbreak attacks against llms
J. Chu, Y. Liu, Z. Yang, X. Shen, M. Backes, and Y. Zhang · 2024
Cited alongside, same era.
Masterkey: Automated jailbreaking of large language model chatbots
Codechameleon: Personalized encryption framework for jailbreaking large language models
H. Lv, X. Wang, Y. Zhang, C. Huang, S. Dou, J. Ye, T. Gui, Q. Zhang, and X. Huang · 2024
Closest in time.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
X. Qi, Y. Zeng, T. Xie, P.-Y. Chen, R. Jia, P. Mittal, and P. Henderson · 2024
Closest in time.
Multilingual blending: Llm safety alignment evaluation with language mixture
J. Song, Y. Huang, Z. Zhou, and L. Ma · 2024
Closest in time.
Reinforcement learning-driven llm agent for automated attacks on llms
X. Wang, J. Peng, K. Xu, H. Yao, and T. Chen · 2024
Closest in time.
Fuzzllm: A novel and universal fuzzing framework for proactively discovering jailbreak vulnerabilities in large language models
D. Yao, J. Zhang, I. G. Harris, and M. Carlsson · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Deng, Y. Liu, Y. Li, K. Wang, Y. Zhang, Z. Li, H. Wang, T. Zhang, and Y. Liu · 2024
Cited alongside, same era.
A wolf in sheep’s clothing: Generalized nested jailbreak prompts can fool large language models easily
P. Ding, J. Kuang, D. Ma, X. Cao, Y. Xian, J. Chen, and S. Huang · 2024
Cited alongside, same era.
Z. Liao and H. Sun · 2024
Cited alongside, same era.
The llama 3 herd of models, 2024
A. . M. Llama Team · 2024
Cited alongside, same era.
When llm meets drl: Advancing jailbreaking efficiency via drl-guided search
X. Chen, Y. Nie, W. Guo, and X. Zhang
Cited in the paper.
Rl-jack: Reinforcement learning-powered black-box jailbreaking attack against llms
X. Chen, Y. Nie, L. Yan, Y. Mao, W. Guo, and X. Zhang
Cited in the paper.
H. Jin, L. Hu, X. Li, P. Zhang, C. Chen, J. Zhuang, and H. Wang
Cited in the paper.
Attackeval: How to evaluate the effectiveness of jailbreak attacking on large language models
M. Jin, S. Zhu, B. Wang, Z. Zhou, C. Zhang, Y. Zhang, et al
Cited in the paper.
Closest in time.
GPT-4 is too smart to be safe: Stealthy chat with LLMs via cipher
Y. Yuan, W. Jiao, W. Wang, J. tse Huang, P. He, S. Shi, and Z. Tu · 2024
Closest in time.
Z. Zhao, X. Zhang, K. Xu, X. Hu, R. Zhang, Z. Du, Q. Guo, and Y. Chen · 2024
Closest in time.
AutoDAN: Interpretable gradient-based adversarial attacks on large language models
S. Zhu, R. Zhang, B. An, G. Wu, J. Barrow, Z. Wang, F. Huang, A. Nenkova, and T. Sun · 2024
Closest in time.