Fetching the paper…
Reading the bibliography…
Growing applications of large language models (LLMs) trained by a third party raise serious concerns on the security vulnerability of LLMs.It has been demonstrated that malicious actors can covertly exploit these vulnerabilities in LLMs through poisoning attacks aimed at generating undesirable outputs.
Bleu: A method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Latent backdoor attacks on deep neural networks
Yuanshun Yao, Huiying Li, Haitao Zheng, and Ben Y. Zhao · 2019
Earlier work this paper cites.
Weight poisoning attacks on pretrained models
Keita Kurita, Paul Michel, and Graham Neubig · 2020
Earlier work this paper cites.
Weight poisoning attacks on pre-trained models, 2020
Keita Kurita, Paul Michel, and Graham Neubig · 2020
Earlier work this paper cites.
Turn the combination lock: Learnable textual backdoor attacks via word substitution
Fanchao Qi, Yuan Yao, Sophia Xu, Zhiyuan Liu, and Maosong Sun · 2021
Earlier work this paper cites.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning, 2021
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Earlier work this paper cites.
Trojaning language models for fun and profit
Xinyang Zhang, Zheng Zhang, Shouling Ji, and Ting Wang · 2021
Earlier work this paper cites.
Badpre: Task-agnostic backdoor attacks to pre-trained nlp foundation models, 2021
Kangjie Chen, Yuxian Meng, Xiaofei Sun, Shangwei Guo, Tianwei Zhang, Jiwei Li, and Chun Fan · 2021
Cited alongside, same era.
Backdoor learning: A survey
Yiming Li, Yong Jiang, Zhifeng Li, and Shu-Tao Xia · 2022
Cited alongside, same era.
Promptattack: Prompt-based attack for language models via gradient search, 2022
Yundi Shi, Piji Li, Changchun Yin, Zhaoyang Han, Lu Zhou, and Zhe Liu · 2022
Cited alongside, same era.
Exploring the universal vulnerability of prompt-based learning paradigm
Lei Xu, Yangyi Chen, Ganqu Cui, Hongcheng Gao, and Zhiyuan Liu · 2022
Cited alongside, same era.
Defending against backdoor attacks in natural language generation, 2022
Xiaofei Sun, Xiaoya Li, Yuxian Meng, Xiang Ao, Lingjuan Lyu, Jiwei Li, and Tianwei Zhang · 2022
Cited alongside, same era.
A survey of natural language generation
Triggerless backdoor attack for nlp tasks with clean labels, 2022
Leilei Gan, Jiwei Li, Tianwei Zhang, Xiaoya Li, Yuxian Meng, Fei Wu, Yi Yang, Shangwei Guo, and Chun Fan · 2022
Later among the works it cites.
Poisoning web-scale training datasets is practical, 2023
Nicholas Carlini, Matthew Jagielski, Christopher A. Choquette-Choo, Daniel Paleka, Will Pearce, Hyrum Anderson, Andreas Terzis, Kurt Thomas, and Florian Tramèr · 2023
Closest in time.
Prompt as triggers for backdoor attack: Examining the vulnerability in language models, 2023
Shuai Zhao, Jinming Wen, Luu Anh Tuan, Junbo Zhao, and Jie Fu · 2023
Closest in time.
Badgpt: Exploring security vulnerabilities of chatgpt via backdoor attacks to instructgpt, 2023
Jiawen Shi, Yixin Liu, Pan Zhou, and Lichao Sun · 2023
Closest in time.
Instructions as backdoors: Backdoor vulnerabilities of instruction tuning for large language models, 2023
Jiashu Xu, Mingyu Derek Ma, Fei Wang, Chaowei Xiao, and Muhao Chen · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chenhe Dong, Yinghui Li, Haifan Gong, Miaoxin Chen, Junxin Li, Ying Shen, and Min Yang · 2022
Cited alongside, same era.
Badprompt: Backdoor attacks on continuous prompts, 2022
Xiangrui Cai, Haidong Xu, Sihan Xu, Ying Zhang, and Xiaojie Yuan · 2022
Cited alongside, same era.
Ppt: Backdoor attacks on pre-trained models via poisoned prompt tuning
Wei Du, Yichun Zhao, Boqun Li, Gongshen Liu, and Shilin Wang · 2022
Cited alongside, same era.
Two-in-one: A model hijacking attack against text generation models, 2023
Wai Man Si, Michael Backes, Yang Zhang, and Ahmed Salem · 2023
Closest in time.
Universal and transferable adversarial attacks on aligned language models, 2023
Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson · 2023
Closest in time.