Fetching the paper…
Reading the bibliography…
Large Language Models face security threats from jailbreak attacks.
2019
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Q. Zheng, X. Xia, X. Zou, Y. Dong, S. Wang, Y. Xue, L. Shen, Z. Wang, A. Wang, Y. Li et al. , “Codegeex: A pre-trained model for code generation with multilingual benchmarking on humaneval-x,” in Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , 2023, pp. 5673–5684
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
2024
Cited alongside, same era.
X. Shen, Z. Chen, M. Backes, Y. Shen, and Y. Zhang, “” do anything now”: Characterizing and evaluating in-the-wild jailbreak prompts on large language models,” in Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security , 2024, pp. 1671–1685
2024
Cited alongside, same era.
G. Deng, Y. Liu, Y. Li, K. Wang, Y. Zhang, Z. Li, H. Wang, T. Zhang, and Y. Liu, “Masterkey: Automated jailbreaking of large language model chatbots,” in NDSS , 2024
2024
Cited alongside, same era.
Z. Yu, X. Liu, S. Liang, Z. Cameron, C. Xiao, and N. Zhang, “Don’t listen to me: Understanding and exploring jailbreak prompts of large language models,” in 33rd USENIX Security Symposium (USENIX Security 24) . Philadelphia, PA: USENIX Association, Aug. 2024, pp. 4675–4692
2024
Later among the works it cites.
2024
Later among the works it cites.
D. Kang, X. Li, I. Stoica, C. Guestrin, M. Zaharia, and T. Hashimoto, “Exploiting programmatic behavior of llms: Dual-use through standard security attacks,” in 2024 IEEE Security and Privacy Workshops (SPW) . IEEE, 2024, pp. 132–143
2024
Later among the works it cites.
2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
F. Jiang, Z. Xu, L. Niu, Z. Xiang, B. Ramasubramanian, B. Li, and R. Poovendran, “Artprompt: Ascii art-based jailbreak attacks against aligned llms,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2024, pp. 15 157–15 173
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Later among the works it cites.
W.-L. Chiang, L. Zheng, Y. Sheng, A. N. Angelopoulos, T. Li, D. Li, B. Zhu, H. Zhang, M. Jordan, J. E. Gonzalez, and I. Stoica, “Chatbot arena: An open platform for evaluating LLMs by human preference,” in Proceedings of the 41st International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, R. Salakhutdinov, Z. Kolter, K. Heller, A. Weller, N. Oliver, J. Scarlett, and F. Berkenkamp, Eds., vol. 235. PMLR, 21–27 Jul 2024, pp. 8359–8388. [Online]. Available: https://proceedings.mlr.press/v235/chiang24b.html
2024
Later among the works it cites.
Anthropic, “Claude 3.7 sonnet system card,” Anthropic, System Card, 2025, Accessed: July 21, 2025. [Online]. Available: https://assets.anthropic.com/m/785e231869ea8b3b/original/claude-3-7-sonnet-system-card.pdf
2025
Closest in time.
user4262, “Dan is my new friend,” Reddit, 2022, Accessed: July 21, 2025. [Online]. Available: https://old.reddit.com/r/ChatGPT/comments/zlcyr9/dan_is_my_new_friend/
2025
Closest in time.
Anthropic, “Prefill Claude’s response for greater output control,” https://web.archive.org/web/20250221204158/https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/prefill-claudes-response
2025
Closest in time.
DeepSeek, “Chat Prefix Completion (Beta),” https://web.archive.org/web/20250307193013/https://api-docs.deepseek.com/guides/chat_prefix_completion
2025
Closest in time.
Google, “Google’s model cards,” https://web.archive.org/web/20250307193013/https://modelcards.withgoogle.com/model-cards
2025
Closest in time.
OpenRouter, “Openrouter,” OpenRouter, 2025, Accessed: July 21, 2025. [Online]. Available: https://openrouter.ai/
2025
Closest in time.