Fetching the paper…
Reading the bibliography…
Universal goal hijacking is a kind of prompt injection attack that forces LLMs to return a target malicious response for arbitrary normal user prompts.
S. Lloyd, “Least squares quantization in pcm,” IEEE transactions on information theory , vol. 28, no. 2, pp. 129–137, 1982
1982
Earlier work this paper cites.
2014
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
N. Carlini and D. Wagner, “Towards evaluating the robustness of neural networks,” in 2017 ieee symposium on security and privacy (sp) . Ieee, 2017, pp. 39–57
2017
Earlier work this paper cites.
E. Wallace, S. Feng, N. Kandpal, M. Gardner, and S. Singh, “Universal adversarial triggers for attacking and analyzing NLP,” in Empirical Methods in Natural Language Processing , 2019
2019
Earlier work this paper cites.
R. Wang, F. Juefei-Xu, Q. Guo, Y. Huang, X. Xie, L. Ma, and Y. Liu, “Amora: Black-box adversarial morphing attack,” in Proceedings of the 28th ACM international conference on multimedia , 2020, pp. 1376–1385
2020
Earlier work this paper cites.
T. Shin, Y. Razeghi, R. L. Logan IV, E. Wallace, and S. Singh, “AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , B. Webber, T. Cohn, Y. He, and Y. Liu, Eds. Online: Association for Computational Linguistics, Nov. 2020, pp. 4222–4235. [Online]. Available: https://aclanthology.org/2020.emnlp-main.346/
2020
Earlier work this paper cites.
Y. Huang, Q. Guo, F. Juefei-Xu, L. Ma, W. Miao, Y. Liu, and G. Pu, “Advfilter: predictive perturbation-aware filtering against adversarial attack via multi-domain learning,” in Proceedings of the 29th ACM International Conference on Multimedia , 2021, pp. 395–403
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
V. Liu and L. B. Chilton, “Design guidelines for prompt engineering text-to-image generative models,” in Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems , 2022, pp. 1–23
2022
Earlier work this paper cites.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou et al. , “Chain-of-thought prompting elicits reasoning in large language models,” Advances in neural information processing systems , vol. 35, pp. 24 824–24 837, 2022
2022
Earlier work this paper cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training language models to follow instructions with human feedback,” Advances in neural information processing systems , vol. 35, pp. 27 730–27 744, 2022
2022
Earlier work this paper cites.
J. Wei, M. Bosma, V. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le, “Finetuned language models are zero-shot learners,” in International Conference on Learning Representations , 2022. [Online]. Available: https://openreview.net/forum?id=gEZrGCozdqR
2022
Earlier work this paper cites.
M. Hao, H. Li, H. Chen, P. Xing, G. Xu, and T. Zhang, “Iron: Private inference on transformers,” Advances in neural information processing systems , vol. 35, pp. 15 718–15 731, 2022
2022
Earlier work this paper cites.
K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, “Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection,” in Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security , 2023, pp. 79–90
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
OpenAI, “Gpt-4,” 2023. [Online]. Available: https://openai.com/research/gpt-4
2023
Cited alongside, same era.
ali, “Qwen,” 2023. [Online]. Available: https://github.com/QwenLM/Qwen
2023
Later among the works it cites.
2023
Later among the works it cites.
andyzoujm, “Advbench,” https://github.com/llm-attacks/llm-attacks/tree/main/data/advbench , 2023
2023
Later among the works it cites.
toughdata, “normal prompt for qqp,” https://huggingface.co/datasets/toughdata/quora-question-answer-dataset , 2023
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
Y. Wen, N. Jain, J. Kirchenbauer, M. Goldblum, J. Geiping, and T. Goldstein, “Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery,” Advances in Neural Information Processing Systems , vol. 36, pp. 51 008–51 025, 2023
2023
Cited alongside, same era.
N. Carlini, M. Nasr, C. A. Choquette-Choo, M. Jagielski, I. Gao, P. W. W. Koh, D. Ippolito, F. Tramer, and L. Schmidt, “Are aligned neural networks adversarially aligned?” Advances in Neural Information Processing Systems , vol. 36, pp. 61 478–61 500, 2023
2023
Cited alongside, same era.
A. Zou, Z. Wang, J. Z. Kolter, and M. Fredrikson, “Universal and transferable adversarial attacks on aligned language models,” 2023
2023
Cited alongside, same era.
F. Shi, X. Chen, K. Misra, N. Scales, D. Dohan, E. H. Chi, N. Schärli, and D. Zhou, “Large language models can be easily distracted by irrelevant context,” in International Conference on Machine Learning . PMLR, 2023, pp. 31 210–31 227
2023
Cited alongside, same era.
2023
Cited alongside, same era.
tastsu lab, “normal prompt for alpaca,” https://huggingface.co/datasets/tatsu-lab/alpaca , 2023
2023
Cited alongside, same era.
Meta, “Llama-2-7b-chat-hf,” https://huggingface.co/meta-llama/Llama-2-7b-chat-hf/ , 2023
2023
Cited alongside, same era.
2024
Closest in time.
Google, “Gemini,” 2024. [Online]. Available: https://gemini.google.com/
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Y. Huang, Q. Guo, F. Juefei-Xu, M. Hu, X. Jia, X. Cao, G. Pu, and Y. Liu, “Texture re-scalable universal adversarial perturbation,” IEEE Transactions on Information Forensics and Security , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Liao and H. Sun, “AmpleGCG: Learning a universal and transferable generative model of adversarial suffixes for jailbreaking both open and closed LLMs,” in First Conference on Language Modeling , 2024. [Online]. Available: https://openreview.net/forum?id=UfqzXg95I5
2024
Closest in time.
A. Köpf, Y. Kilcher, D. von Rütte, S. Anagnostidis, Z. R. Tam, K. Stevens, A. Barhoum, D. Nguyen, O. Stanley, R. Nagyfi et al. , “Openassistant conversations-democratizing large language model alignment,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
X. Zhang, C. Zhang, T. Li, Y. Huang, X. Jia, M. Hu, J. Zhang, Y. Liu, S. Ma, and C. Shen, “Jailguard: A universal detection framework for prompt-based attacks on llm systems,” ACM Trans. Softw. Eng. Methodol. , Mar. 2025, just Accepted. [Online]. Available: https://doi.org/10.1145/3724393
2025
Closest in time.
Y. Huang, X. Luo, Q. Guo, F. Juefei-Xu, X. Jia, W. Miao, G. Pu, and Y. Liu, “Scale-invariant adversarial attack against arbitrary-scale super-resolution,” IEEE Transactions on Information Forensics and Security , 2025
2025
Closest in time.
Y. Huang, L. Liang, T. Li, X. Jia, R. Wang, W. Miao, G. Pu, and Y. Liu, “Perception-guided jailbreak against text-to-image models,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 39, no. 25, 2025, pp. 26 238–26 247
2025
Closest in time.