The Large Language Models (LLMs) are poised to offer efficient and intelligent services for future mobile communication networks, owing to their exceptional capabilities in language comprehension and generation.
However, the extremely high data and computational resource requirements for the performance of LLMs compel developers to resort to outsourcing training or utilizing third-party data and computing resources.
These strategies may expose the model within the network to maliciously manipulated training data and processing, providing an opportunity for attackers to embed a hidden backdoor into the model, termed a backdoor attack.
Backdoor attack in LLMs refers to embedding a hidden backdoor in LLMs that causes the model to perform normally on benign samples but exhibit degraded performance on poisoned ones.
A Comprehensive Overview of Backdoor Attacks in Large Language Models within Communication Networks · Around
Li S, Liu H, Dong T, et al. Hidden backdoors in human-centric language models[C]//Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security. 2021: 3123-3140
2021
Earlier work this paper cites.
Zhang X, Zhang Z, Ji S, et al. Trojaning language models for fun and profit[C]//2021 IEEE European Symposium on Security and Privacy (EuroS&P). IEEE, 2021: 179-197
Pan X, Zhang M, Sheng B, et al. Hidden trigger backdoor attack on NLP models via linguistic style manipulation[C]//31st USENIX Security Symposium (USENIX Security 22). 2022: 3611-3628
2022
Cited alongside, same era.
Cai X, Xu H, Xu S, et al. Badprompt: Backdoor attacks on continuous prompts[J]. Advances in Neural Information Processing Systems, 2022, 35: 37068-37080
Lyu W, Zheng S, Ling H, et al. Backdoor Attacks Against Transformers with Attention Enhancement[C]//ICLR 2023 Workshop on Backdoor Attacks and Defenses in Machine Learning. 2023