Fetching the paper…
Reading the bibliography…
Large language models (LLMs) and LLM-based agents have been widely deployed in a wide range of applications in the real world, including healthcare diagnostics, financial analysis, customer support, robotics, and autonomous driving, expanding their powerful capability of understanding, reasoning, and generating natural languages.
1910
Earlier work this paper cites.
D. Lee and M. Yannakakis, “Principles and methods of testing finite state machines-a survey,” Proceedings of the IEEE , vol. 84, no. 8, pp. 1090–1123, 1996
1996
Earlier work this paper cites.
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers) , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
D. J. Fremont, T. Dreossi, S. Ghosh, X. Yue, A. L. Sangiovanni-Vincentelli, and S. A. Seshia, “Scenic: A language for scenario specification and scene generation,” in Proceedings of the 40th ACM SIGPLAN conference on programming language design and implementation , 2019, pp. 63–78
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
S. Li, H. Liu, T. Dong, B. Z. H. Zhao, M. Xue, H. Zhu, and J. Lu, “Hidden Backdoors in Human-Centric Language Models,” in Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security , 2021, pp. 3123–3140
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
OpenAI, “ChatGPT (Feb 20 Version),” 2023. [Online]. Available: https://openai.com/chatgpt
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
J. Xue, M. Zheng, T. Hua, Y. Shen, Y. Liu, L. Bölöni, and Q. Lou, “TrojLLM: A Black-box Trojan Prompt Attack on Large Language Models,” Advances in Neural Information Processing Systems , vol. 36, pp. 65 665–65 677, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
N. Gu, P. Fu, X. Liu, Z. Liu, Z. Lin, and W. Wang, “A Gradient Control Method for Backdoor Attacks on Parameter-Efficient Tuning,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2023, pp. 3508–3520
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Wei, N. Haghtalab, and J. Steinhardt, “Jailbroken: How Does LLM Safety Training Fail?” Advances in Neural Information Processing Systems , vol. 36, pp. 80 079–80 110, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Mehrotra, M. Zampetakis, P. Kassianik, B. Nelson, H. Anderson, Y. Singer, and A. Karbasi, “Tree of Attacks: Jailbreaking Black-Box LLMs Automatically,” 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Mendes, “Ultimate ChatGPT Prompt Engineering Guide for General Users and Developers,” 2023. [Online]. Available: https://www.imaginarycloud.com/blog/chatgpt-prompt-engineering
2023
Cited alongside, same era.
“Sandwich defense,” 2023. [Online]. Available: {https://learnprompting.org/docs/prompt\_hacking/defensive\_measures/\\sandwich\_defense}
2023
Cited alongside, same era.
“Instruction defense,” 2023. [Online]. Available: \url{https://learnprompting.org/docs/prompt\_hacking/defensive\_measures/\\instruction}
2023
Cited alongside, same era.
2023
Cited alongside, same era.
E. Yudkowsky, “Using gpt: Eliezer against chatgpt jailbreaking,” 2023. [Online]. Available: \url{https://www.alignmentforum.org/posts/pNcFYZnPdXyL2RfgA/using-gpt-eliezer-against-chatgpt-jailbreaking}
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Zeng, H. Lin, J. Zhang, D. Yang, R. Jia, and W. Shi, “How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2024, pp. 14 322–14 350
2024
Later among the works it cites.
Y. Liu, Y. Jia, R. Geng, J. Jia, and N. Z. Gong, “Formalizing and Benchmarking Prompt Injection Attacks and Defenses,” in 33rd USENIX Security Symposium (USENIX Security 24) , 2024, pp. 1831–1847
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
K. Krishna, Y. Song, M. Karpinska, J. Wieting, and M. Iyyer, “Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense,” Advances in Neural Information Processing Systems , vol. 36, pp. 27 469–27 500, 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
J. Li, T. Tang, W. X. Zhao, J.-Y. Nie, and J.-R. Wen, “Pre-Trained Language Models for Text Generation: A Survey,” ACM Computing Surveys , vol. 56, no. 9, pp. 1–39, 2024
2024
Cited alongside, same era.
2024
Later among the works it cites.
W. Zhang, X. Kong, C. Dewitt, T. Braunl, and J. B. Hong, “A Study on Prompt Injection Attack Against LLM-Integrated Mobile Robotic Systems,” in 2024 IEEE 35th International Symposium on Software Reliability Engineering Workshops (ISSREW) , 2024, pp. 361–368
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
J. Shi, Z. Yuan, Y. Liu, Y. Huang, P. Zhou, L. Sun, and N. Z. Gong, “Optimization-based Prompt Injection Attack to LLM-as-a-Judge,” in Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security , 2024, pp. 660–674
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
M. L. Siddiq, J. Zhang, and J. C. D. S. Santos, “Understanding Regular Expression Denial of Service (ReDoS): Insights from LLM-Generated Regexes and Developer Forums,” in Proceedings of the 32nd IEEE/ACM International Conference on Program Comprehension , 2024, pp. 190–201
2024
Later among the works it cites.
Z. Shi, Y. Wang, F. Yin, X. Chen, K.-W. Chang, and C.-J. Hsieh, “Red Teaming Language Model Detectors with Language Models,” Transactions of the Association for Computational Linguistics , vol. 12, pp. 174–189, 2024
2024
Later among the works it cites.
M. Christ, S. Gunn, and O. Zamir, “Undetectable Watermarks for Language Models,” in The Thirty Seventh Annual Conference on Learning Theory . PMLR, 2024, pp. 1125–1139
2024
Later among the works it cites.
2024
Later among the works it cites.
2025
Closest in time.
xAI, “Grok-3,” 2025, [Large language model]. [Online]. Available: https://grok.com/
2025
Closest in time.
S. Zhao, M. Jia, Z. Guo, L. Gan, X. Xu, X. Wu, J. Fu, F. Yichao, F. Pan, and A. T. Luu, “A Survey of Recent Backdoor Attacks and Defenses in Large Language Models,” Transactions on Machine Learning Research , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
B. C. Das, M. H. Amini, and Y. Wu, “Security and Privacy Challenges of Large Language Models: A Survey,” ACM Comput. Surv. , vol. 57, no. 6, Feb. 2025. [Online]. Available: https://doi.org/10.1145/3712001
2025
Closest in time.
Learn Prompting, “Prompt Hacking: Jailbreaking,” 2025, accessed: 2025-03-01. [Online]. Available: https://learnprompting.org/docs/prompt\_hacking/jailbreaking\#footnotes
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Y. Xue, J. Wang, Z. Yin, Y. Ma, H. Qin, R. Tao, and X. Liu, “Dual Intention Escape: Jailbreak Attack against Large Language Models,” in THE WEB CONFERENCE 2025 , 2025
2025
Closest in time.
S. Willison, “Delimiters Won’t Save You,” https://simonwillison.net/2023/May/11/delimiters-wont-save-you/
2025
Closest in time.