Fetching the paper…
Reading the bibliography…
Jailbreak attacks represent one of the most sophisticated threats to the security of large language models (LLMs).
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training language models to follow instructions with human feedback,” Advances in Neural Information Processing Systems , 2022
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Robey, E. Wong, H. Hassani, and G. J. Pappas, “Smoothllm: Defending large language models against jailbreaking attacks,” 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
B. Wang, W. Chen, and H. P. et al., “Decodingtrust: A comprehensive assessment of trustworthiness in gpt models,” 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
X. Liu, N. Xu, M. Chen, and C. Xiao, “Autodan: Generating stealthy jailbreak prompts on aligned large language models,” 2023
2023
Earlier work this paper cites.
Z. Wang, F. Yang, L. Wang, P. Zhao, H. Wang, L. Chen, Q. Lin, and K.-F. Wong, “Self-guard: Empower the llm to safeguard itself,” 2023
2023
Earlier work this paper cites.
Y. Liu, G. Deng, Z. Xu, Y. Li, Y. Zheng, Y. Zhang, L. Zhao, T. Zhang, and Y. Liu, “Jailbreaking chatgpt via prompt engineering: An empirical study,” 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
A. Albert. (2023) Jailbreak chat. https://www.jailbreakchat.com/
2023
Cited alongside, same era.
2024
Cited alongside, same era.
Z. Zhou, Q. Wang, M. Jin, J. Yao, J. Ye, W. Liu, W. Wang, X. Huang, and K. Huang, “Mathattack: Attacking large language models towards math solving ability,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 17, 2024, pp. 19 750–19 758
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Y. Yao, J. Duan, K. Xu, Y. Cai, Z. Sun, and Y. Zhang, “A survey on large language model (llm) security and privacy: The good, the bad, and the ugly,” High-Confidence Computing , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y. Wang et al. , “A survey on evaluation of large language models,” ACM Transactions on Intelligent Systems and Technology , 2024
2024
Cited alongside, same era.
M. Jin, Q. Yu, D. Shu, H. Zhao, W. Hua, Y. Meng, Y. Zhang, and M. Du, “The impact of reasoning step length on large language models,” in Findings of the Association for Computational Linguistics ACL 2024 , Bangkok, Thailand and virtual meeting, Aug. 2024, pp. 1830–1842. [Online]. Available: https://aclanthology.org/2024.findings-acl.108
2024
Cited alongside, same era.
Z. Li, B. Peng, P. He, M. Galley, J. Gao, and X. Yan, “Guiding large language models via directional stimulus prompting,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
C. Zhang, M. Jin, Q. Yu, C. Liu, H. Xue, and X. Jin, “Goal-guided generative prompt injection attack on large language models,” in 2024 IEEE International Conference on Data Mining (ICDM) , 2024, pp. 941–946
2024
Closest in time.
C. Zhang, M. Jin, D. Shu, T. Wang, D. Liu, and X. Jin, “Target-driven attack for large language models,” in ECAI , 2024
2024
Closest in time.
2024
Closest in time.
T. Wang, Z. Fang, H. Xue, C. Zhang, M. Jin, W. Xu, D. Shu, S. Yang, Z. Wang, and D. Liu, “Large vision-language model security: A survey,” in Frontiers in Cyber Security , B. Chen, X. Fu, and M. Huang, Eds. Singapore: Springer Nature Singapore, 2024, pp. 3–22
2024
Closest in time.
2024
Closest in time.
Reddit contributors, “Chatgptjailbreak subreddit,” https://www.reddit.com/r/ChatGPTJailbreak/ , 2024
2024
Closest in time.