Fetching the paper…
Reading the bibliography…
Prompt injection has emerged as a serious security threat to large language models (LLMs).
MPNet: Masked and Permuted Pre-training for Language Understanding
Song, K.; Tan, X.; Qin, T.; Lu, J.; and Liu, T.-Y. 2020 · 2004
Earlier work this paper cites.
Learning and Classification of Malware Behavior
Rieck, K.; Holz, T.; Willems, C.; Düssel, P.; and Laskov, P. 2008 · 2008
Earlier work this paper cites.
Extracting Training Data from Large Language Models
Carlini, N.; Tramer, F.; Wallace, E.; Jagielski, M.; Herbert-Voss, A.; Lee, K.; Roberts, A.; Brown, T.; Song, D.; Erlingsson, U.; Oprea, A.; and Raffel, C. 2021 · 2012
Earlier work this paper cites.
AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
Shin, T.; Razeghi, Y.; au2, R. L. L. I.; Wallace, E.; and Singh, S. 2020 · 2020
Earlier work this paper cites.
Bad Characters: Imperceptible NLP Attacks
Boucher, N.; Shumailov, I.; Anderson, R.; and Papernot, N. 2021 · 2021
Earlier work this paper cites.
BadNL: Backdoor Attacks against NLP Models with Semantic-preserving Improvements
Chen, X.; Salem, A.; Chen, D.; Backes, M.; Ma, S.; Shen, Q.; Wu, Z.; and Zhang, Y. 2021 · 2021
Earlier work this paper cites.
Quantifying Privacy Risks of Masked Language Models Using Membership Inference Attacks
Mireshghallah, F.; Goyal, K.; Uniyal, A.; Berg-Kirkpatrick, T.; and Shokri, R. 2022 · 2022
Earlier work this paper cites.
Ignore Previous Prompt: Attack Techniques For Language Models
Perez, F.; and Ribeiro, I. 2022 · 2022
Earlier work this paper cites.
Black-Box Tuning for Language-Model-as-a-Service
Sun, T.; Shao, Y.; Qian, H.; Huang, X.; and Qiu, X. 2022 · 2022
Cited alongside, same era.
How is ChatGPT’s behavior changing over time?
Chen, L.; Zaharia, M.; and Zou, J. 2023 · 2023
Cited alongside, same era.
More than You’ve Asked for: A Comprehensive Analysis of Novel Prompt Injection Threats to Application-Integrated Large Language Models
Greshake, K.; Abdelnabi, S.; Mishra, S.; Endres, C.; Holz, T.; and Fritz, M. 2023 · 2023
Cited alongside, same era.
Multi-step Jailbreaking Privacy Attacks on ChatGPT
Li, H.; Guo, D.; Fan, W.; Xu, M.; Huang, J.; Meng, F.; and Song, Y. 2023 · 2023
Cited alongside, same era.
Prompt Injection attack against LLM-integrated Applications
Liu, Y.; Deng, G.; Li, Y.; Wang, K.; Zhang, T.; Liu, Y.; Wang, H.; Zheng, Y.; and Liu, Y. 2023 · 2023
Cited alongside, same era.
Analyzing Leakage of Personally Identifiable Information in Language Models
r/ChatGPTJailbreak
Reddit. 2023 · 2023
Closest in time.
”Do Anything Now”: Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
Shen, X.; Chen, Z.; Backes, M.; Shen, Y.; and Zhang, Y. 2023 · 2023
Closest in time.
Safeguarding Crowdsourcing Surveys from ChatGPT with Prompt Injection
Wang, C.; Freire, S. K.; Zhang, M.; Wei, J.; Goncalves, J.; Kostakos, V.; Sarsenbayeva, Z.; Schneegass, C.; Bozzon, A.; and Niforatos, E. 2023 · 2023
Closest in time.
Wen, R.; Wang, T.; Backes, M.; Zhang, Y.; and Salem, A. 2023 · 2023
Closest in time.
Defending ChatGPT against Jailbreak Attack via Self-Reminder
Wu, F.; Xie, Y.; Yi, J.; Shao, J.; Curl, J.; Lyu, L.; Chen, Q.; and Xie, X. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lukas, N.; Salem, A.; Sim, R.; Tople, S.; Wutschitz, L.; and Zanella-Béguelin, S. 2023 · 2023
Cited alongside, same era.
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
Mehrotra, A.; Zampetakis, M.; Kassianik, P.; Nelson, B.; Anderson, H.; Singer, Y.; and Karbasi, A. 2023 · 2023
Cited alongside, same era.
Automatic Prompt Optimization with ”Gradient Descent” and Beam Search
Pryzant, R.; Iter, D.; Li, J.; Lee, Y. T.; Zhu, C.; and Zeng, M. 2023 · 2023
Cited alongside, same era.
Backdoor Attacks for In-Context Learning with Language Models
Kandpal, N.; Jagielski, M.; Tramèr, F.; and Carlini, N. 2023a
Cited in the paper.
User Inference Attacks on Large Language Models
Kandpal, N.; Pillutla, K.; Oprea, A.; Kairouz, P.; Choquette-Choo, C. A.; and Xu, Z. 2023b
Cited in the paper.
Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection
Yan, J.; Yadav, V.; Li, S.; Chen, L.; Tang, Z.; Wang, H.; Srinivasan, V.; Ren, X.; and Jin, H. 2023 · 2023
Closest in time.
GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher
Yuan, Y.; Jiao, W.; Wang, W.; tse Huang, J.; He, P.; Shi, S.; and Tu, Z. 2023 · 2023
Closest in time.
Universal and Transferable Adversarial Attacks on Aligned Language Models
Zou, A.; Wang, Z.; Kolter, J. Z.; and Fredrikson, M. 2023 · 2023
Closest in time.