Fetching the paper…
Reading the bibliography…
LLMs are commonly used in retrieval-augmented applications to execute user instructions based on data from external sources.
L. Van der Maaten and G. Hinton, “Visualizing data using t-sne.” Journal of machine learning research , vol. 9, no. 11, 2008
2008
Earlier work this paper cites.
F. Schroff, D. Kalenichenko, and J. Philbin, “Facenet: A unified embedding for face recognition and clustering,” in CVPR , 2015
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in ICLR , 2015
2015
Earlier work this paper cites.
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang, “Squad: 100,000+ questions for machine comprehension of text,” in EMNLP , 2016
2016
Earlier work this paper cites.
Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. Cohen, R. Salakhutdinov, and C. D. Manning, “Hotpotqa: A dataset for diverse, explainable multi-hop question answering,” in EMNLP , 2018
2018
Earlier work this paper cites.
F. Perez and I. Ribeiro, “Ignore previous prompt: Attack techniques for language models,” in NeurIPS ML Safety Workshop , 2022
2022
Earlier work this paper cites.
C. Burns, H. Ye, D. Klein, and J. Steinhardt, “Discovering latent knowledge in language models without supervision,” in ICLR , 2022
2022
Earlier work this paper cites.
K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, “Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection,” in ACM CCS AISec Workshop , 2023
2023
Earlier work this paper cites.
J. Yi, Y. Xie, B. Zhu, K. Hines, E. Kiciman, G. Sun, X. Xie, and F. Wu, “Benchmarking and defending against indirect prompt injection attacks on large language models,” arXiv , 2023
2023
Earlier work this paper cites.
A. Zou, Z. Wang, J. Z. Kolter, and M. Fredrikson, “Universal and transferable adversarial attacks on aligned language models,” arXiv , 2023
2023
Earlier work this paper cites.
A. Zou, L. Phan, S. Chen, J. Campbell, P. Guo, R. Ren, A. Pan, X. Yin, M. Mazeika, A.-K. Dombrowski et al. , “Representation engineering: A top-down approach to ai transparency,” arXiv , 2023
2023
Earlier work this paper cites.
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. Hashimoto, “Alpaca: a strong, replicable instruction-following model,” 2023
2023
Earlier work this paper cites.
S. Chaudhary, “Code alpaca: An instruction-following llama model for code generation,” https://github.com/sahil280114/codealpaca , 2023
2023
Earlier work this paper cites.
J. Ji, M. Liu, J. Dai, X. Pan, C. Zhang, C. Bian, B. Chen, R. Sun, Y. Wang, and Y. Yang, “Beavertails: Towards improved safety alignment of llm via a human-preference dataset,” in NeurIPS Datasets and Benchmarks Track , 2023
2023
Earlier work this paper cites.
P. Chao, A. Robey, E. Dobriban, H. Hassani, G. J. Pappas, and E. Wong, “Jailbreaking black box large language models in twenty queries,” arXiv , 2023
2023
Earlier work this paper cites.
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier et al. , “Mistral 7b,” arXiv , 2023
2023
Earlier work this paper cites.
Z. X. Yong, C. Menghini, and S. Bach, “Low-resource languages jailbreak gpt-4,” in NeurIPS Workshop of Socially Responsible Language Modelling Research , 2023
2023
Earlier work this paper cites.
R. Hendel, M. Geva, and A. Globerson, “In-context learning creates task vectors,” in Findings of EMNLP , 2023
2023
Cited alongside, same era.
L. Zhang, R. T. McCoy, T. R. Sumers, J.-Q. Zhu, and T. L. Griffiths, “Deep de finetti: Recovering topic distributions from large language models,” arXiv , 2023
2023
Cited alongside, same era.
J. Chu, Y. Liu, Z. Yang, X. Shen, M. Backes, and Y. Zhang, “Comprehensive assessment of jailbreak attacks against LLMs,” arXiv , 2024
2024
Cited alongside, same era.
K. Hines, G. Lopez, M. Hall, F. Zarfati, Y. Zunger, and E. Kiciman, “Defending against indirect prompt injection attacks with spotlighting,” arXiv , 2024
2024
Cited alongside, same era.
J. Piet, M. Alrashed, C. Sitawarin, S. Chen, Z. Wei, E. Sun, B. Alomair, and D. Wagner, “Jatmo: Prompt injection defense by task-specific finetuning,” in European Symposium on Research in Computer Security , 2024
Y. Zeng, W. Sun, T. Huynh, D. Song, B. Li, and R. Jia, “Beear: Embedding-based adversarial removal of safety backdoors in instruction-tuned language models,” in EMNLP , 2024
2024
Closest in time.
N. Belrose, Z. Furman, L. Smith, D. Halawi, I. Ostrovsky, L. McKinney, S. Biderman, and J. Steinhardt, “Eliciting latent predictions from transformers with the tuned lens,” arXiv , 2024
2024
Closest in time.
P. Ding, J. Kuang, D. Ma, X. Cao, Y. Xian, J. Chen, and S. Huang, “A wolf in sheep’s clothing: Generalized nested jailbreak prompts can fool large language models easily,” in NAACL , 2024
2024
Closest in time.
X. Shen, Z. Chen, M. Backes, Y. Shen, and Y. Zhang, ““Do Anything Now”: Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models,” in ACM CCS , 2024
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
A. Vassilev, A. Oprea, A. Fordyce, and H. Anderson, “Adversarial machine learning: A taxonomy and terminology of attacks and mitigations,” National Institute of Standards and Technology, Tech. Rep., 2024
2024
Cited alongside, same era.
J. Ferrando, G. Sarti, A. Bisazza, and M. R. Costa-jussà, “A primer on the inner workings of transformer-based language models,” arXiv , 2024
2024
Cited alongside, same era.
D. Halawi, J.-S. Denain, and J. Steinhardt, “Overthinking the truth: Understanding how language models process false demonstrations,” in ICLR , 2024
2024
Cited alongside, same era.
A. T. Mallen, M. Brumley, J. Kharchenko, and N. Belrose, “Eliciting latent knowledge from” quirky” language models,” in COLM , 2024
2024
Cited alongside, same era.
Microsoft, “Prompt shields,” https://learn.microsoft.com/en-us/azure/ai-services/content-safety/concepts/jailbreak-detection , 2024
2024
Cited alongside, same era.
A. RoyChowdhury, M. Luo, P. Sahu, S. Banerjee, and M. Tiwari, “Confusedpilot: Compromising enterprise information integrity and confidentiality with copilot for microsoft 365,” arXiv , 2024
2024
Cited alongside, same era.
D. Pasquini, M. Strohmeier, and C. Troncoso, “Neural exec: Learning (and learning from) execution triggers for prompt injection attacks,” in ACM CCS AISec Workshop , 2024
2024
Cited alongside, same era.
Y. Huang, L. Sun, H. Wang, S. Wu, Q. Zhang, Y. Li, C. Gao, Y. Huang, W. Lyu, Y. Zhang, X. Li, H. Sun, Z. Liu, Y. Liu, Y. Wang, Z. Zhang, B. Vidgen, B. Kailkhura, C. Xiong, C. Xiao, C. Li, E. P. Xing, F. Huang, H. Liu, H. Ji, H. Wang, H. Zhang, H. Yao, M. Kellis, M. Zitnik, M. Jiang, M. Bansal, J. Zou, J. Pei, J. Liu, J. Gao, J. Han, J. Zhao, J. Tang, J. Wang, J. Vanschoren, J. Mitchell, K. Shu, K. Xu, K.-W. Chang, L. He, L. Huang, M. Backes, N. Z. Gong, P. S. Yu, P.-Y. Chen, Q. Gu, R. Xu, R. Ying, S. Ji, S. Jana, T. Chen, T. Liu, T. Zhou, W. Y. Wang, X. Li, X. Zhang, X. Wang, X. Xie, X. Chen, X. Wang, Y. Liu, Y. Ye, Y. Cao, Y. Chen, and Y. Zhao, “Trustllm: Trustworthiness in large language models,” in ICML , 2024
2024
Closest in time.
Y. Wang, H. Li, X. Han, P. Nakov, and T. Baldwin, “Do-not-answer: Evaluating safeguards in LLMs,” in Findings of EACL , 2024
2024
Closest in time.
P. Chao, E. Debenedetti, A. Robey, M. Andriushchenko, F. Croce, V. Sehwag, E. Dobriban, N. Flammarion, G. J. Pappas, F. Tramèr, H. Hassani, and E. Wong, “Jailbreakbench: An open robustness benchmark for jailbreaking large language models,” in NeurIPS Datasets and Benchmarks , 2024
2024
Closest in time.
M. Abdin, S. A. Jacobs, A. A. Awan, J. Aneja, A. Awadallah, H. Awadalla, N. Bach, A. Bahree, A. Bakhtiari, H. Behl et al. , “Phi-3 technical report: A highly capable language model locally on your phone,” arXiv , 2024
2024
Closest in time.
Meta, “Introducing meta Llama 3: The most capable openly available LLM to date,” https://ai.meta.com/blog/meta-llama-3/ , 2024
2024
Closest in time.
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. d. l. Casas, E. B. Hanna, F. Bressand et al. , “Mixtral of experts,” arXiv , 2024
2024
Closest in time.
W. Zhao, X. Ren, J. Hessel, C. Cardie, Y. Choi, and Y. Deng, “WildChat: 1m ChatGPT interaction logs in the wild,” in ICLR , 2024
2024
Closest in time.
E. Debenedetti, J. Rando, D. Paleka, S. F. Florin, D. Albastroiu, N. Cohen, Y. Lemberg, R. Ghosh, R. Wen, A. Salem et al. , “Dataset and lessons learned from the 2024 SaTML LLM Capture-the-Flag competition,” in NeurIPS Dataset and Benchmarks , 2024
2024
Closest in time.
A. Fay, S. Abdelnabi, B. Pannell, G. Cherubin, A. Salem, A. Paverd, C. M. Amhlaoibh, J. Rakita, S. Zanella-Beguelin, E. Zverev, M. Russinovich, and J. Rando3, “Llmail-inject: Adaptive prompt injection challenge,” https://microsoft.github.io/llmail-inject/ , 2024
2024
Closest in time.
E. Zverev, S. Abdelnabi, M. Fritz, and C. H. Lampert, “Can llms separate instructions from data? and what do we even mean by that?” in ICLR , 2025
2025
Closest in time.
S. Chen, J. Piet, C. Sitawarin, and D. Wagner, “Struq: Defending against prompt injection with structured queries,” in USENIX Security , 2025
2025
Closest in time.
W. Zou, R. Geng, B. Wang, and J. Jia, “Poisonedrag: Knowledge corruption attacks to retrieval-augmented generation of large language models,” in USENIX Security , 2025
2025
Closest in time.
M. Andriushchenko, F. Croce, and N. Flammarion, “Jailbreaking leading safety-aligned llms with simple adaptive attacks,” in ICLR , 2025
2025
Closest in time.