Fetching the paper…
Reading the bibliography…
Prompt injection attack, where an attacker injects a prompt into the original one, aiming to make an Large Language Model (LLM) follow the injected prompt to perform an attacker-chosen task, represent a critical security threat.
ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out . Association for Computational Linguistics, Barcelona, Spain, 74–81
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Automatically Constructing a Corpus of Sentential Paraphrases. In IWP
William B. Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
Contributions to the Study of SMS Spam Filtering: New Collection and Results. In DOCENG
Tiago A. Almeida, Jose Maria Gomez Hidalgo, and Akebo Yamakami. 2011 · 2011
Earlier work this paper cites.
Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank. In EMNLP
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Predicting Grammaticality on an Ordinal Scale. In ACL
Michael Heilman, Aoife Cahill, Nitin Madnani, Melissa Lopez, Matthew Mulholland, and Joel Tetreault. 2014 · 2014
Earlier work this paper cites.
A Neural Attention Model for Abstractive Sentence Summarization
Alexander M. Rush, Sumit Chopra, and Jason Weston. 2015 · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences. In NeurIPS
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Earlier work this paper cites.
Automated Hate Speech Detection and the Problem of Offensive Language. In ICWSM
Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017 · 2017
Earlier work this paper cites.
JFLEG: A Fluency Corpus and Benchmark for Grammatical Error Correction. In EACL
Courtney Napoles, Keisuke Sakaguchi, and Joel Tetreault. 2017 · 2017
Earlier work this paper cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. In ICLR
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2019 · 2019
Earlier work this paper cites.
Towards crowdsourced training of large neural networks using decentralized mixture-of-experts. In NeurIPS
Max Ryabinin and Anton Gusev. 2020 · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding. In ICLR
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2021 · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback. In NeurIPS
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Ignore Previous Prompt: Attack Techniques For Language Models. In NeurIPS ML Safety Workshop
Fábio Perez and Ian Ribeiro. 2022 · 2022
Cited alongside, same era.
Prompt injection attacks against GPT-3
Simon Willison. 2022 · 2022
Cited alongside, same era.
The falcon series of open language models
Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Mérouane Debbah, Étienne Goffinet, Daniel Hesslow, Julien Launay, Quentin Malartic, et al · 2023
Cited alongside, same era.
How Amazon continues to improve the customer reviews experience with generative AI
Amazon. 2023 · 2023
Cited alongside, same era.
Not what you’ve signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. In AISec
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. 2023 · 2023
Cited alongside, same era.
Orca DPO Pairs
Intel. 2024 · 2024
Closest in time.
Automatic and universal prompt injection attacks against large language models
Xiaogeng Liu, Zhiyuan Yu, Yizhe Zhang, Ning Zhang, and Chaowei Xiao. 2024b · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al · 2024
Closest in time.
Neural Exec: Learning (and Learning from) Execution Triggers for Prompt Injection Attacks
Dario Pasquini, Martin Strohmeier, and Carmela Troncoso. 2024 · 2024
Closest in time.
Is poisoning a real threat to LLM alignment? Maybe more so than you think
Pankayaraj Pathmanathan, Souradip Chakraborty, Xiangyu Liu, Yongyuan Liang, and Furong Huang. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rich Harang. 2023 · 2023
Cited alongside, same era.
OWASP Top 10 for Large Language Model Applications
OWASP. 2023 · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model. In NeurIPS
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023 · 2023
Cited alongside, same era.
Universal jailbreak backdoors from poisoned human feedback
Javier Rando and Florian Tramèr. 2023 · 2023
Cited alongside, same era.
LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 · 2023
Cited alongside, same era.
Delimiters won’t save you from prompt injection
Simon Willison. 2023 · 2023
Cited alongside, same era.
Universal and Transferable Adversarial Attacks on Aligned Language Models
Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson. 2023 · 2023
Cited alongside, same era.
Gpqa: A graduate-level google-proof q&a benchmark. In COLM
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman. 2024 · 2024
Closest in time.
BAIT: Large Language Model Backdoor Scanning by Inverting Attack Target. In IEEE S & P
Guangyu Shen, Siyuan Cheng, Zhuo Zhang, Guanhong Tao, Kaiyuan Zhang, Hanxi Guo, Lu Yan, Xiaolong Jin, Shengwei An, Shiqing Ma, et al · 2024
Closest in time.
Optimization-based Prompt Injection Attack to LLM-as-a-Judge. In ACM CCS
Jiawen Shi, Zenghui Yuan, Yinuo Liu, Yue Huang, Pan Zhou, Lichao Sun, and Neil Zhenqiang Gong. 2024 · 2024
Closest in time.
CodecLM: Aligning Language Models with Tailored Synthetic Data
Zifeng Wang, Chun-Liang Li, Vincent Perot, Long T Le, Jin Miao, Zizhao Zhang, Chen-Yu Lee, and Tomas Pfister. 2024a · 2024
Closest in time.
Preference Poisoning Attacks on Reward Model Learning
Junlin Wu, Jiongxiao Wang, Chaowei Xiao, Chenguang Wang, Ning Zhang, and Yevgeniy Vorobeychik. 2024 · 2024
Closest in time.
Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing
Zhangchen Xu, Fengqing Jiang, Luyao Niu, Yuntian Deng, Radha Poovendran, Yejin Choi, and Bill Yuchen Lin. 2024a · 2024
Closest in time.
A Critical Evaluation of Defenses against Prompt Injection Attacks
Yuqi Jia, Zedian Shao, Yupei Liu, Jinyuan Jia, Dawn Song, and Neil Zhenqiang Gong. 2025 · 2025
Closest in time.
DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks. In IEEE Symposium on Security and Privacy
Yupei Liu, Yuqi Jia, Jinyuan Jia, Dawn Song, and Neil Zhenqiang Gong. 2025 · 2025
Closest in time.
Prompt Injection Attack to Tool Selection in LLM Agents
Jiawen Shi, Zenghui Yuan, Guiyao Tie, Pan Zhou, Neil Zhenqiang Gong, and Lichao Sun. 2025 · 2025
Closest in time.
EnvInjection: Environmental Prompt Injection Attack to Multi-modal Web Agents
Xilong Wang, John Bloch, Zedian Shao, Yuepeng Hu, Shuyan Zhou, and Neil Zhenqiang Gong. 2025 · 2025
Closest in time.
Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models. In ICLR
Biao Yi, Tiansheng Huang, Sishuo Chen, Tong Li, Zheli Liu, Chu Zhixuan, and Yiming Li. 2025 · 2025
Closest in time.