Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are vulnerable to backdoor attacks that manipulate outputs via hidden triggers.
Onion: A simple and effective defense against textual backdoor attacks, 2021a
Qi, F., Chen, Y., Li, M., Yao, Y., Liu, Z., and Sun, M · 2011
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Goodfellow, I. J., Shlens, J., and Szegedy, C · 2015
Earlier work this paper cites.
Learning both weights and connections for efficient neural networks
Han, S., Pool, J., Tran, J., and Dally, W. J · 2015
Earlier work this paper cites.
Parseval networks: Improving robustness to adversarial examples
Cisse, M., Bojanowski, P., Grave, E., Dauphin, Y., and Usunier, N · 2017
Earlier work this paper cites.
Fine-pruning: Defending against backdooring attacks on deep neural networks
Liu, K., Dolan-Gavitt, B., and Garg, S · 2018
Earlier work this paper cites.
A backdoor attack against lstm-based text classification systems
Dai, J., Chen, C., and Li, Y · 2019
Earlier work this paper cites.
Badnets: Evaluating backdooring attacks on deep neural networks
Gu, T., Liu, K., Dolan-Gavitt, B., and Garg, S · 2019
Earlier work this paper cites.
Qusecnets: Quantization-based defense mechanism for securing deep neural network against adversarial attacks
Khalid, F., Ali, H., Tariq, H., Hanif, M. A., Rehman, S., Ahmed, R., and Shafique, M · 2019
Earlier work this paper cites.
Latent backdoor attacks on deep neural networks
Yao, Y., Li, H., Zheng, H., and Zhao, B. Y · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y · 2020
Earlier work this paper cites.
Weight poisoning attacks on pretrained models
Kurita, K., Michel, P., and Neubig, G · 2020
Earlier work this paper cites.
Neural attention distillation: Erasing backdoor triggers from deep neural networks
Li, Y., Lyu, X., Koren, N., Lyu, L., Li, B., and Ma, X · 2021
Earlier work this paper cites.
Hidden killer: Invisible textual backdoor attacks with syntactic trigger
Qi, F., Li, M., Chen, Y., Zhang, Z., Liu, Z., Wang, Y., and Sun, M · 2021
Earlier work this paper cites.
Adversarial robustness with semi-infinite constrained learning
Robey, A., Chamon, L., Pappas, G. J., Hassani, H., and Ribeiro, A · 2021
Earlier work this paper cites.
Concealed data poisoning attacks on NLP models
Wallace, E., Zhao, T., Feng, S., and Singh, S · 2021
Earlier work this paper cites.
Adversarial neuron pruning purifies backdoored deep models
Wu, D. and Wang, Y · 2021
Earlier work this paper cites.
RAP: Robustness-Aware Perturbations for defending against backdoor attacks on NLP models
Yang, W., Lin, Y., Li, P., Zhou, J., and Sun, X · 2021
Earlier work this paper cites.
A robustly optimized BERT pre-training approach with post-training
Zhuang, L., Wayne, L., Ya, S., and Jun, Z · 2021
Earlier work this paper cites.
Universal and transferable adversarial attacks on aligned language models, 2023
Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., and Fredrikson, M · 2021
Earlier work this paper cites.
Spinning language models: Risks of propaganda-as-a-service and countermeasures
Bagdasaryan, E. and Shmatikov, V · 2022
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback, 2022
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., Chen, C., Olsson, C., Olah, C., Hernandez, D., Drain, D., Ganguli, D., Li, D., Tran-Johnson, E., Perez, E., Kerr, J., Mueller, J., Ladish, J., Landau, J., Ndousse, K., Lukosuite, K., Lovitt, L., Sellitto, M., Elhage, N., Schiefer, N., Mercado, N., DasSarma, N., Lasenby, R., Larson, R., Ringer, S., Johnston, S., Kravec, S., Showk, S. E., Fort, S., Lanham, T., Telleen-Lawton, T., Conerly, T., Henighan, T., Hume, T., Bowman, S. R., Hatfield-Dodds, Z., Mann, B., Amodei, D., Joseph, N., McCandlish, S., Brown, T., and Kaplan, J · 2022
Cited alongside, same era.
Kallima: A clean-label framework for textual backdoor attacks
Chen, X., Dong, Y., Sun, Z., Zhai, S., Shen, Q., and Wu, Z · 2022
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Cited alongside, same era.
Constrained optimization with dynamic bound-scaling for effective NLP backdoor defense
Shen, G., Liu, Y., Tao, G., Xu, Q., Zhang, Z., An, S., Ma, S., and Zhang, X · 2022
Cited alongside, same era.
BITE: Textual backdoor attacks with iterative trigger injection
Yan, J., Gupta, V., and Ren, X · 2023
Later among the works it cites.
Prompt as triggers for backdoor attack: Examining the vulnerability in language models
Zhao, S., Wen, J., Luu, A., Zhao, J., and Fu, J · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I · 2023
Later among the works it cites.
Adversarially Guided Stateful Defense Against Backdoor Attacks in Federated Deep Learning
Ali, H., Nepal, S., Kanhere, S. S., and Jha, S · 2024
Closest in time.
Poisoning Web-Scale Training Datasets is Practical
Carlini, N., Jagielski, M., Choquette-Choo, C. A., Paleka, D., Pearce, W., Anderson, H., Terzis, A., Thomas, K., and Tramer, F · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2022
Cited alongside, same era.
Pre-activation distributions expose backdoor neurons
Yue, Z., Liu, Z., Sun, H., Zhang, G., Zhang, J., Zhang, J., Li, J., and Liu, H · 2022
Cited alongside, same era.
Fine-mixing: Mitigating backdoors in fine-tuned language models
Zhang, Z., Lyu, L., Ma, X., Wang, C., and Sun, X · 2022
Cited alongside, same era.
Data-free backdoor removal based on channel lipschitzness
Zheng, R., Tang, R., Li, J., and Liu, L · 2022
Cited alongside, same era.
Backdoor learning on sequence to sequence models, 2023
Chen, L., Cheng, M., and Huang, H · 2023
Cited alongside, same era.
On the effectiveness of adversarial training against backdoor attacks
Gao, Y., Wu, D., Zhang, J., Gan, G., Xia, S.-T., Niu, G., and Sugiyama, M · 2023
Cited alongside, same era.
Defending against backdoor attacks by layer-wise feature analysis
Jebreel, N. M., Domingo-Ferrer, J., and Li, Y · 2023
Cited alongside, same era.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Cited alongside, same era.
Chuang, Y., Xie, Y., Luo, H., Kim, Y., Glass, J. R., and He, P · 2024
Closest in time.
Composite backdoor attacks against large language models
Huang, H., Zhao, Z., Backes, M., Shen, Y., and Zhang, Y · 2024
Closest in time.
Sleeper agents: Training deceptive llms that persist through safety training
Hubinger, E., Denison, C., Mu, J., Lambert, M., Tong, M., MacDiarmid, M., Lanham, T., Ziegler, D. M., Maxwell, T., Cheng, N., Jermyn, A. S., Askell, A., Radhakrishnan, A., Anil, C., Duvenaud, D., Ganguli, D., Barez, F., Clark, J., Ndousse, K., Sachan, K., Sellitto, M., Sharma, M., DasSarma, N., Grosse, R. B., Kravec, S., Bai, Y., Witten, Z., Favaro, M., Brauner, J., Karnofsky, H., Christiano, P. F., Bowman, S. R., Graham, L., Kaplan, J., Mindermann, S., Greenblatt, R., Shlegeris, B., Schiefer, N., and Perez, E · 2024
Closest in time.
Defending against backdoor attacks by layer-wise feature analysis (extended abstract)
Jebreel, N. M., Domingo-Ferrer, J., and Li, Y · 2024
Closest in time.
CleanGen: Mitigating backdoor attacks for generation tasks in large language models
Li, Y., Xu, Z., Jiang, F., Niu, L., Sahabandu, D., Ramasubramanian, B., and Poovendran, R · 2024
Closest in time.
Crow: Implementation for llm backdoor elimination, 2024
Min, N. M · 2024
Closest in time.
Unified neural backdoor removal with only few clean samples through unlearning and relearning, 2024
Min, N. M., Pham, L. H., and Sun, J · 2024
Closest in time.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Qi, X., Zeng, Y., Xie, T., Chen, P.-Y., Jia, R., Mittal, P., and Henderson, P · 2024
Closest in time.
Code llama: Open foundation models for code, 2024
Rozière, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Sauvestre, R., Remez, T., Rapin, J., Kozhevnikov, A., Evtimov, I., Bitton, J., Bhatt, M., Ferrer, C. C., Grattafiori, A., Xiong, W., Défossez, A., Copet, J., Azhar, F., Touvron, H., Martin, L., Usunier, N., Scialom, T., and Synnaeve, G · 2024
Closest in time.
Instructions as backdoors: Backdoor vulnerabilities of instruction tuning for large language models
Xu, J., Ma, M., Wang, F., Xiao, C., and Chen, M · 2024
Closest in time.
Backdooring instruction-tuned large language models with virtual prompt injection
Yan, J., Yadav, V., Li, S., Chen, L., Tang, Z., Wang, H., Srinivasan, V., Ren, X., and Jin, H · 2024
Closest in time.
Instruction backdoor attacks against customized LLMs
Zhang, R., Li, H., Wen, R., Jiang, W., Zhang, Y., Backes, M., Shen, Y., and Zhang, Y · 2024
Closest in time.
Unlearning backdoor attacks for llms with weak-to-strong knowledge distillation, 2024
Zhao, S., Wu, X., Nguyen, C.-D., Jia, M., Feng, Y., and Tuan, L. A · 2024
Closest in time.
Tuba: Cross-lingual transferability of backdoor attacks in llms with instruction tuning, 2025
He, X., Wang, J., Xu, Q., Minervini, P., Stenetorp, P., Rubinstein, B. I. P., and Cohn, T · 2025
Closest in time.
Li, Y., Huang, H., Zhao, Y., Ma, X., and Sun, J · 2025
Closest in time.