Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are vulnerable to attacks like prompt injection, backdoor attacks, and adversarial attacks, which manipulate prompts or models to generate harmful outputs.
Targeted backdoor attacks on deep learning systems using data poisoning
Xinyun Chen, Chang Liu, Bo Li, Kimberly Lu, and Dawn Song. 2017 · 2017
Earlier work this paper cites.
Adversarial attacks on neural network policies
Sandy H. Huang, Nicolas Papernot, Ian J. Goodfellow, Yan Duan, and Pieter Abbeel. 2017 · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. 2017 · 2017
Earlier work this paper cites.
Threat of adversarial attacks on deep learning in computer vision: A survey
Naveed Akhtar and Ajmal S. Mian. 2018 · 2018
Earlier work this paper cites.
Smooth loss functions for deep top-k classification
Leonard Berrada, Andrew Zisserman, and M. Pawan Kumar. 2018 · 2018
Earlier work this paper cites.
Boosting adversarial attacks with momentum
Yinpeng Dong, Fangzhou Liao, Tianyu Pang, Hang Su, Jun Zhu, Xiaolin Hu, and Jianguo Li. 2018 · 2018
Earlier work this paper cites.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein. 2018 · 2018
Earlier work this paper cites.
Smoothness and stability in gans
Casey Chu, Kentaro Minami, and Kenji Fukumizu. 2020 · 2020
Earlier work this paper cites.
Hidden trigger backdoor attacks
Aniruddha Saha, Akshayvarun Subramanya, and Hamed Pirsiavash. 2020 · 2020
Earlier work this paper cites.
Gradient-based adversarial attacks against text transformers
Chuan Guo, Alexandre Sablayrolles, Hervé Jégou, and Douwe Kiela. 2021 · 2021
Earlier work this paper cites.
ONION: A simple and effective defense against textual backdoor attacks
Fanchao Qi, Yangyi Chen, Mukai Li, Yuan Yao, Zhiyuan Liu, and Maosong Sun. 2021 · 2021
Earlier work this paper cites.
RAP: robustness-aware perturbations for defending against backdoor attacks on NLP models
Wenkai Yang, Yankai Lin, Peng Li, Jie Zhou, and Xu Sun. 2021 · 2021
Earlier work this paper cites.
Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection
Sahar Abdelnabi, Kai Greshake, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. 2023 · 2023
Earlier work this paper cites.
Detecting language model attacks with perplexity
Gabriel Alon and Michael Kamfonas. 2023 · 2023
Earlier work this paper cites.
Revisit input perturbation problems for llms: A unified robustness evaluation framework for noisy slot filling task
Guanting Dong, Jinxu Zhao, Tingfeng Hui, Daichi Guo, Wenlong Wang, Boqi Feng, Yueyan Qiu, Zhuoma Gongque, Keqing He, Zechen Wang, and Weiran Xu. 2023 · 2023
Earlier work this paper cites.
Large language models for code: Security hardening and adversarial testing
Jingxuan He and Martin T. Vechev. 2023 · 2023
Earlier work this paper cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and Madian Khabsa. 2023 · 2023
Earlier work this paper cites.
Baseline defenses for adversarial attacks against aligned language models
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping-yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein. 2023 · 2023
Earlier work this paper cites.
Backdoor attacks for in-context learning with language models
Nikhil Kandpal, Matthew Jagielski, Florian Tramèr, and Nicholas Carlini. 2023 · 2023
Cited alongside, same era.
Certifying LLM safety against adversarial prompting
Aounon Kumar, Chirag Agarwal, Suraj Srinivas, Soheil Feizi, and Hima Lakkaraju. 2023 · 2023
Cited alongside, same era.
Machine unlearning in gradient boosting decision trees
Huawei Lin, Jun Woo Chung, Yingjie Lao, and Weijie Zhao. 2023 · 2023
Cited alongside, same era.
Prompt injection attack against llm-integrated applications
Yi Liu, Gelei Deng, Yuekang Li, Kailong Wang, Tianwei Zhang, Yepang Liu, Haoyu Wang, Yan Zheng, and Yang Liu. 2023 · 2023
Cited alongside, same era.
Mitigating over-smoothing in transformers via regularized nonlocal functionals
Tam Nguyen, Tan Nguyen, and Richard G. Baraniuk. 2023 · 2023
Adversarial attacks and defenses for large language models (llms): methods, frameworks & challenges
Pranjal Kumar. 2024 · 2024
Later among the works it cites.
Formalizing and benchmarking prompt injection attacks and defenses
Yupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia, and Neil Zhenqiang Gong. 2024 · 2024
Later among the works it cites.
Inkit Padhi, Manish Nagireddy, Giandomenico Cornacchia, Subhajit Chaudhury, Tejaswini Pedapati, Pierre L. Dognin, Keerthiram Murugesan, Erik Miehling, Martin Santillan Cooper, Kieran Fraser, Giulio Zizzo, Muhammad Zaid Hameed, Mark Purcell, Michael Desmond, Qian Pan, Zahra Ashktorab, Inge Vejsbjerg, Elizabeth M. Daly, Michael Hind, Werner Geyer, Ambrish Rawat, Kush R. Varshney, and Prasanna Sattigeri. 2024 · 2024
Later among the works it cites.
Jatmo: Prompt injection defense by task-specific finetuning
Julien Piet, Maha Alrashed, Chawin Sitawarin, Sizhe Chen, Zeming Wei, Elizabeth Sun, Basel Alomair, and David A. Wagner. 2024 · 2024
Later among the works it cites.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Prompting large language models with answer heuristics for knowledge-based visual question answering
Zhenwei Shao, Zhou Yu, Meng Wang, and Jun Yu. 2023 · 2023
Cited alongside, same era.
Survey of vulnerabilities in large language models revealed by adversarial attacks
Erfan Shayegani, Md Abdullah Al Mamun, Yu Fu, Pedram Zaree, Yue Dong, and Nael B. Abu-Ghazaleh. 2023 · 2023
Cited alongside, same era.
Alpaca: A strong, replicable instruction-following model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023 · 2023
Cited alongside, same era.
Are large language models really robust to word-level perturbations?
Haoyu Wang, Guozheng Ma, Cong Yu, Ning Gui, Linrui Zhang, Zhiqi Huang, Suwei Ma, Yongzhe Chang, Sen Zhang, Li Shen, Xueqian Wang, Peilin Zhao, and Dacheng Tao. 2023 · 2023
Cited alongside, same era.
A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly
Yifan Yao, Jinhao Duan, Kaidi Xu, Yuanfang Cai, Eric Sun, and Yue Zhang. 2023 · 2023
Cited alongside, same era.
Prompting large language model for machine translation: A case study
Biao Zhang, Barry Haddow, and Alexandra Birch. 2023 · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson. 2023 · 2023
Cited alongside, same era.
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. 2024 · 2024
Later among the works it cites.
Is llm-as-a-judge robust? investigating universal adversarial attacks on zero-shot LLM assessment
Vyas Raina, Adian Liusie, and Mark J. F. Gales. 2024 · 2024
Later among the works it cites.
Lightweight safety classification using pruned language models
Mason Sawtell, Tula Masterman, Sandi Besen, and Jim Brown. 2024 · 2024
Later among the works it cites.
Optimization-based prompt injection attack to llm-as-a-judge
Jiawen Shi, Zenghui Yuan, Yinuo Liu, Yue Huang, Pan Zhou, Lichao Sun, and Neil Zhenqiang Gong. 2024 · 2024
Later among the works it cites.
Badagent: Inserting and activating backdoor attacks in LLM agents
Yifei Wang, Dizhan Xue, Shengjie Zhang, and Shengsheng Qian. 2024 · 2024
Later among the works it cites.
A new era in LLM security: Exploring security concerns in real-world llm-based systems
Fangzhou Wu, Ning Zhang, Somesh Jha, Patrick D. McDaniel, and Chaowei Xiao. 2024 · 2024
Later among the works it cites.
Badchain: Backdoor chain-of-thought prompting for large language models
Zhen Xiang, Fengqing Jiang, Zidi Xiong, Bhaskar Ramasubramanian, Radha Poovendran, and Bo Li. 2024 · 2024
Later among the works it cites.
An LLM can fool itself: A prompt-based adversarial attack
Xilie Xu, Keyi Kong, Ning Liu, Lizhen Cui, Di Wang, Jingfeng Zhang, and Mohan S. Kankanhalli. 2024 · 2024
Later among the works it cites.
Backdooring instruction-tuned large language models with virtual prompt injection
Jun Yan, Vikas Yadav, Shiyang Li, Lichang Chen, Zheng Tang, Hai Wang, Vijay Srinivasan, Xiang Ren, and Hongxia Jin. 2024 · 2024
Later among the works it cites.
Watch out for your agents! investigating backdoor threats to llm-based agents
Wenkai Yang, Xiaohan Bi, Yankai Lin, Sishuo Chen, Jie Zhou, and Xu Sun. 2024 · 2024
Later among the works it cites.
Universal vulnerabilities in large language models: Backdoor attacks for in-context learning
Shuai Zhao, Meihuizi Jia, Anh Tuan Luu, Fengjun Pan, and Jinming Wen. 2024 · 2024
Later among the works it cites.
On prompt-driven safeguarding for large language models
Chujie Zheng, Fan Yin, Hao Zhou, Fandong Meng, Jie Zhou, Kai-Wei Chang, Minlie Huang, and Nanyun Peng. 2024 · 2024
Later among the works it cites.
Adversarial attacks on large language models
Jing Zou, Shungeng Zhang, and Meikang Qiu. 2024 · 2024
Later among the works it cites.