Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have revolutionized various domains but remain vulnerable to prompt injection attacks, where malicious inputs manipulate the model into ignoring original instructions and executing designated action.
Backdoor attacks and countermeasures on deep learning: A comprehensive review
Yansong Gao, Bao Gia Doan, Zhi Zhang, Siqi Ma, Jiliang Zhang, Anmin Fu, Surya Nepal, and Hyoungshick Kim. 2020 · 2007
Earlier work this paper cites.
Hidden trigger backdoor attacks
Aniruddha Saha, Akshayvarun Subramanya, and Hamed Pirsiavash. 2020 · 2020
Earlier work this paper cites.
Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2021 · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. 2021 · 2021
Earlier work this paper cites.
A study of the attention abnormality in trojaned BERTs
Weimin Lyu, Songzhu Zheng, Tengfei Ma, and Chao Chen. 2022 · 2022
Earlier work this paper cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al. 2022 · 2022
Earlier work this paper cites.
Ignore previous prompt: Attack techniques for language models
Fábio Perez and Ian Ribeiro. 2022 · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Earlier work this paper cites.
Detecting language model attacks with perplexity
Gabriel Alon and Michael Kamfonas. 2023 · 2023
Earlier work this paper cites.
Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. 2023 · 2023
Earlier work this paper cites.
Baseline defenses for adversarial attacks against aligned language models
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping-yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein. 2023 · 2023
Earlier work this paper cites.
Prompt injection attack against llm-integrated applications
Yi Liu, Gelei Deng, Yuekang Li, Kailong Wang, Zihao Wang, Xiaofeng Wang, Tianwei Zhang, Yepang Liu, Haoyu Wang, Yan Zheng, et al. 2023 · 2023
Earlier work this paper cites.
Learn Prompting: Your Guide to Communicating with AI — learnprompting.org
2023 · 2024
Earlier work this paper cites.
Are you still on track!? catching llm task drift with activations
Sahar Abdelnabi, Aideen Fay, Giovanni Cherubin, Ahmed Salem, Mario Fritz, and Andrew Paverd. 2024 · 2024
Earlier work this paper cites.
Phi-3 technical report: A highly capable language model locally on your phone
Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl, et al. 2024 · 2024
Earlier work this paper cites.
Struq: Defending against prompt injection with structured queries
Sizhe Chen, Julien Piet, Chawin Sitawarin, and David Wagner. 2024 · 2024
Cited alongside, same era.
Yung-Sung Chuang, Linlu Qiu, Cheng-Yu Hsieh, Ranjay Krishna, Yoon Kim, and James Glass. 2024 · 2024
Cited alongside, same era.
Induction heads as an essential mechanism for pattern matching in in-context learning
J. Crosbie and E. Shutova. 2024 · 2024
Cited alongside, same era.
Dataset and lessons learned from the 2024 satml llm capture-the-flag competition
Edoardo Debenedetti, Javier Rando, Daniel Paleka, Silaghi Fineas Florin, Dragos Albastroiu, Niv Cohen, Yuval Lemberg, Reshmi Ghosh, Rui Wen, Ahmed Salem, et al. 2024 · 2024
Cited alongside, same era.
GitHub - protectai/rebuff: LLM Prompt Injection Detector — github.com
ProtectAI.com. 2024b · 2024
Closest in time.
Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. 2024 · 2024
Closest in time.
Optimization-based prompt injection attack to llm-as-a-judge
Jiawen Shi, Zenghui Yuan, Yinuo Liu, Yue Huang, Pan Zhou, Lichao Sun, and Neil Zhenqiang Gong. 2024 · 2024
Closest in time.
Rethinking interpretability in the era of large language models
Chandan Singh, Jeevana Priya Inala, Michel Galley, Rich Caruana, and Jianfeng Gao. 2024 · 2024
Closest in time.
Using GPT-Eliezer against ChatGPT Jailbreaking — LessWrong — lesswrong.com
rgorman Stuart Armstrong. 2022 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
deepset/prompt-injections · Datasets at Hugging Face — huggingface.co
deepset. 2023 · 2024
Cited alongside, same era.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024 · 2024
Cited alongside, same era.
A primer on the inner workings of transformer-based language models
Javier Ferrando, Gabriele Sarti, Arianna Bisazza, and Marta R Costa-jussà. 2024 · 2024
Cited alongside, same era.
Successor heads: Recurring, interpretable attention heads in the wild
Rhys Gould, Euan Ong, George Ogden, and Arthur Conmy. 2024 · 2024
Cited alongside, same era.
Defending against indirect prompt injection attacks with spotlighting
Keegan Hines, Gary Lopez, Matthew Hall, Federico Zarfati, Yonatan Zunger, and Emre Kiciman. 2024 · 2024
Cited alongside, same era.
Prompt injection attacks in defended systems
Daniil Khomsky, Narek Maloyan, and Bulat Nutfullin. 2024 · 2024
Cited alongside, same era.
Prompt Guard-86M | Model Cards and Prompt formats — llama.com
Meta. 2024 · 2024
Cited alongside, same era.
Owasp top 10 for llm applications
OWASP. 2023 · 2024
Cited alongside, same era.
Xuchen Suo. 2024 · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al. 2024 · 2024
Closest in time.
Function vectors in large language models
Eric Todd, Millicent Li, Arnab Sen Sharma, Aaron Mueller, Byron C Wallace, and David Bau. 2024 · 2024
Closest in time.
Tensor trust: Interpretable prompt injection attacks from an online game
Sam Toyer, Olivia Watkins, Ethan Adrian Mendes, Justin Svegliato, Luke Bailey, Tiffany Wang, Isaac Ong, Karim Elmaaroufi, Pieter Abbeel, Trevor Darrell, Alan Ritter, and Stuart Russell. 2024 · 2024
Closest in time.
The instruction hierarchy: Training llms to prioritize privileged instructions
Eric Wallace, Kai Xiao, Reimar Leike, Lilian Weng, Johannes Heidecke, and Alex Beutel. 2024 · 2024
Closest in time.
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al. 2024 · 2024
Closest in time.
Poisonprompt: Backdoor attack on prompt-based large language models
Hongwei Yao, Jian Lou, and Zhan Qin. 2024 · 2024
Closest in time.
x.com — x.com
Yohei. 2022 · 2024
Closest in time.
Can llms separate instructions from data? and what do we even mean by that?
Egor Zverev, Sahar Abdelnabi, Mario Fritz, and Christoph H Lampert. 2024 · 2024
Closest in time.