Fetching the paper…
Reading the bibliography…
Large language models (LLMs) that integrate multiple input roles (e.g., system instructions, user queries, external tool outputs) are increasingly prevalent in practice.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Ignore previous prompt: Attack techniques for language models
Perez, F. and Ribeiro, I · 2022
Earlier work this paper cites.
Prompt injection attacks against GPT-3, 2022
Willison, S · 2022
Earlier work this paper cites.
Extending context window of large language models via positional interpolation
Chen, S., Wong, S., Chen, L., and Tian, Y · 2023
Earlier work this paper cites.
Yarn: Efficient context window extension of large language models
Peng, B., Quesnelle, J., Fan, H., and Shippole, E · 2023
Earlier work this paper cites.
Ignore this title and HackAPrompt: Exposing systemic vulnerabilities of llms through a global scale prompt hacking competition
Schulhoff, S., Pinto, J., Khan, A., Bouchard, L.-F., Si, C., Anati, S., Tagliabue, V., Kost, A. L., Carnahan, C., and Boyd-Graber, J · 2023
Earlier work this paper cites.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Cited alongside, same era.
Tensor trust: Interpretable prompt injection attacks from an online game
Toyer, S., Watkins, O., Mendes, E. A., Svegliato, J., Bailey, L., Wang, T., Ong, I., Elmaaroufi, K., Abbeel, P., Darrell, T., et al · 2023
Cited alongside, same era.
Efficient streaming language models with attention sinks
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2023
Cited alongside, same era.
Assessing prompt injection risks in 200+ custom gpts
Yu, J., Wu, Y., Shu, D., Jin, M., and Xing, X · 2023
Cited alongside, same era.
Pose: Efficient context window extension of llms via positional skip-wise training
Struq: Defending against prompt injection with structured queries
Chen, S., Piet, J., Sitawarin, C., and Wagner, D · 2024
Later among the works it cites.
Coercing LLMs to do and reveal (almost) anything
Geiping, J., Stein, A., Shu, M., Saifullah, K., Wen, Y., and Goldstein, T · 2024
Later among the works it cites.
The instruction hierarchy: Training llms to prioritize privileged instructions
Wallace, E., Xiao, K., Leike, R., Weng, L., Heidecke, J., and Beutel, A · 2024
Later among the works it cites.
Instructional segment embedding: Improving llm safety with instruction hierarchy
Wu, T., Zhang, S., Song, K., Xu, S., Zhao, S., Agrawal, R., Indurthi, S. R., Xiang, C., Mittal, P., and Zhou, W · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhu, D., Yang, N., Wang, L., Song, Y., Wu, W., Wei, F., and Li, S · 2023
Cited alongside, same era.
Llama 3 model card
AI@Meta · 2024
Cited alongside, same era.
Gandalf: Ignore instructions
Lakera AI
Cited in the paper.
Gandalf: Summarization
Lakera AI
Cited in the paper.
Yu, J., Shao, Y., Miao, H., Shi, J., and Xing, X · 2024
Later among the works it cites.
Can LLMs separate instructions from data? And what do we even mean by that?
Zverev, E., Abdelnabi, S., Fritz, M., and Lampert, C. H · 2024
Later among the works it cites.