Fetching the paper…
Reading the bibliography…
Language Models (LMs) are becoming increasingly popular in real-world applications.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th annual meeting of the Association for Computational Linguistics
2002
Earlier work this paper cites.
C.-Y. Lin, “Rouge: A package for automatic evaluation of summaries,” in Text summarization branches out
2004
Earlier work this paper cites.
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
D. Cer, Y. Yang, S.-y. Kong, N. Hua, N. Limtiaco, R. S. John, N. Constant, M. Guajardo-Cespedes, S. Yuan, C. Tar, et al
2018
Earlier work this paper cites.
K. Liu, B. Dolan-Gavitt, and S. Garg, “Fine-pruning: Defending against backdooring attacks on deep neural networks,” in International symposium on research in attacks, intrusions, and defenses
2018
Earlier work this paper cites.
G. Siracusano, M. Trevisan, R. Gonzalez, and R. Bifulco, “Poster: on the application of nlp to discover relationships between malicious network entities,” in Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security
2019
Earlier work this paper cites.
J. Dai, C. Chen, and Y. Li, “A backdoor attack against lstm-based text classification systems,” IEEE Access
2019
Earlier work this paper cites.
Y. Gao, C. Xu, D. Wang, S. Chen, D. C. Ranasinghe, and S. Nepal, “Strip: A defence against trojan attacks on deep neural networks,” in Proceedings of the 35th Annual Computer Security Applications Conference
2019
Earlier work this paper cites.
N. Reimers and I. Gurevych, “Sentence-bert: Sentence embeddings using siamese bert-networks,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)
2019
Earlier work this paper cites.
Q. Feng, D. He, Z. Liu, H. Wang, and K.-K. R. Choo, “Securenlp: A system for multi-party privacy-preserving natural language processing,” IEEE Transactions on Information Forensics and Security
2020
Earlier work this paper cites.
M. Zulqarnain, R. Ghazali, Y. M. M. Hassim, and M. Rehan, “A comparative review on deep learning models for text classification,” Indones. J. Electr. Eng. Comput. Sci
2020
Earlier work this paper cites.
Z. Tan, S. Wang, Z. Yang, G. Chen, X. Huang, M. Sun, and Y. Liu, “Neural machine translation: A review of methods, resources, and tools,” AI Open
2020
Earlier work this paper cites.
K. Kurita, P. Michel, and G. Neubig, “Weight poisoning attacks on pretrained models,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
S. Garg, A. Kumar, V. Goel, and Y. Liang, “Can adversarial weight perturbations inject neural backdoors,” in Proceedings of the 29th ACM International Conference on Information & Knowledge Management
2020
Earlier work this paper cites.
Y. Huang, Z. Song, D. Chen, K. Li, and S. Arora, “Texthide: Tackling data privacy in language understanding tasks,” in Findings of the Association for Computational Linguistics: EMNLP 2020
2020
Earlier work this paper cites.
S. Li, H. Liu, T. Dong, B. Z. H. Zhao, M. Xue, H. Zhu, and J. Lu, “Hidden backdoors in human-centric language models,” in Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security
2021
Earlier work this paper cites.
F. Qi, Y. Yao, S. Xu, Z. Liu, and M. Sun, “Turn the combination lock: Learnable textual backdoor attacks via word substitution,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)
2021
Earlier work this paper cites.
F. Qi, M. Li, Y. Chen, Z. Zhang, Z. Liu, Y. Wang, and M. Sun, “Hidden killer: Invisible textual backdoor attacks with syntactic trigger,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)
2021
Earlier work this paper cites.
F. Qi, Y. Chen, X. Zhang, M. Li, Z. Liu, and M. Sun, “Mind the style of text! adversarial and backdoor attacks based on text style transfer,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing
2021
Earlier work this paper cites.
X. Chen, A. Salem, D. Chen, M. Backes, S. Ma, Q. Shen, Z. Wu, and Y. Zhang, “Badnl: Backdoor attacks against nlp models with semantic-preserving improvements,” in 37th Annual Computer Security Applications Conference, ACSAC 2021
2021
Earlier work this paper cites.
L. Shen, S. Ji, X. Zhang, J. Li, J. Chen, J. Shi, C. Fang, J. Yin, and T. Wang, “Backdoor pre-trained models can transfer to all,” in Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security
2021
Earlier work this paper cites.
F. Qi, Y. Chen, M. Li, Y. Yao, Z. Liu, and M. Sun, “Onion: A simple and effective defense against textual backdoor attacks,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
K. Chen, Y. Meng, X. Sun, S. Guo, T. Zhang, J. Li, and C. Fan, “Badpre: Task-agnostic backdoor attacks to pre-trained nlp foundation models,” in International Conference on Learning Representations
2021
Earlier work this paper cites.
C. Xu, J. Wang, Y. Tang, F. Guzmán, B. I. Rubinstein, and T. Cohn, “A targeted attack on black-box neural machine translation with parallel data poisoning,” in Proceedings of the web conference 2021
2021
Earlier work this paper cites.
L. Li, D. Song, X. Li, J. Zeng, R. Ma, and X. Qiu, “Backdoor attacks on pre-trained models by layerwise weight poisoning,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing
2021
Earlier work this paper cites.
W. Yang, L. Li, Z. Zhang, X. Ren, X. Sun, and B. He, “Be careful about poisoned word embeddings: Exploring the vulnerability of the embedding layers in nlp models,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
2021
Earlier work this paper cites.
Z. Zhang, X. Ren, Q. Su, X. Sun, and B. He, “Neural network surgery: Injecting data patterns into pre-trained models with minimal instance-wise side effects,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
2021
Earlier work this paper cites.
H. Kwon and S. Lee, “Textual backdoor attack for the text classification system,” Security and Communication Networks
2021
Earlier work this paper cites.
W. Yang, Y. Lin, P. Li, J. Zhou, and X. Sun, “Rethinking stealthiness of backdoor attack against nlp models,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)
2021
Earlier work this paper cites.
X. C. A. Salem and M. Zhang, “Badnl: Backdoor attacks against nlp models,” in ICML 2021 Workshop on Adversarial Machine Learning
2021
Earlier work this paper cites.
K. Shao, J. Yang, Y. Ai, H. Liu, and Y. Zhang, “Bddr: An effective defense against textual backdoor attacks,” Computers & Security
2021
Earlier work this paper cites.
W. Yang, Y. Lin, P. Li, J. Zhou, and X. Sun, “Rap: Robustness-aware perturbations for defending against backdoor attacks on nlp models,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing
2021
Earlier work this paper cites.
T. L. Le, N. P. Park, and D. Lee, “A sweet rabbit hole by darcy: Using honeypots to detect universal trigger’s adversarial attacks,” in 59th Annual Meeting of the Association for Comp. Linguistics (ACL)
2021
Earlier work this paper cites.
Z. Li, D. Mekala, C. Dong, and J. Shang, “Bfclass: A backdoor-free text classification framework,” in Findings of the Association for Computational Linguistics: EMNLP 2021
2021
Earlier work this paper cites.
M. Fan, Z. Si, X. Xie, Y. Liu, and T. Liu, “Text backdoor detection using an interpretable rnn abstract model,” IEEE Transactions on Information Forensics and Security
2021
Earlier work this paper cites.
C. Chen and J. Dai, “Mitigating backdoor attacks in lstm-based text classification systems by backdoor keyword identification,” Neurocomputing
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
X. Zhang, Z. Zhang, S. Ji, and T. Wang, “Trojaning language models for fun and profit,” in 2021 IEEE European Symposium on Security and Privacy (EuroS&P)
2021
Earlier work this paper cites.
X. Xu, Q. Wang, H. Li, N. Borisov, C. A. Gunter, and B. Li, “Detecting ai trojans using meta neural analysis,” in 2021 IEEE Symposium on Security and Privacy (SP)
2021
Earlier work this paper cites.
E. Wallace, T. Zhao, S. Feng, and S. Singh, “Concealed data poisoning attacks on nlp models,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
2021
Earlier work this paper cites.
J. Wang, C. Xu, F. Guzmán, A. El-Kishky, Y. Tang, B. Rubinstein, and T. Cohn, “Putting words into the system’s mouth: A targeted attack on neural machine translation using monolingual data poisoning,” in Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021
2021
Earlier work this paper cites.
F. Alsharadgah, A. Khreishah, M. Al-Ayyoub, Y. Jararweh, G. Liu, I. Khalil, M. Almutiry, and N. Saeed, “An adaptive black-box defense against trojan attacks on text data,” in 2021 Eighth International Conference on Social Network Analysis, Management and Security (SNAMS)
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
T. H. Alwaneen, A. M. Azmi, H. A. Aboalsamh, E. Cambria, and A. Hussain, “Arabic question answering system: a survey,” Artificial Intelligence Review
2022
Earlier work this paper cites.
S. Li, T. Dong, B. Z. H. Zhao, M. Xue, S. Du, and H. Zhu, “Backdoors against natural language processing: A review,” IEEE Security & Privacy
2022
Earlier work this paper cites.
G. Cui, L. Yuan, B. He, Y. Chen, Z. Liu, and M. Sun, “A unified evaluation of textual backdoor learning: Frameworks and benchmarks,” in Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track
2022
Earlier work this paper cites.
L. Xu, Y. Chen, G. Cui, H. Gao, and Z. Liu, “Exploring the universal vulnerability of prompt-based learning paradigm,” in Findings of the Association for Computational Linguistics: NAACL 2022
2022
Earlier work this paper cites.
W. Du, Y. Zhao, B. Li, G. Liu, and S. Wang, “Ppt: Backdoor attacks on pre-trained models via poisoned prompt tuning,” in Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI-22
2022
Earlier work this paper cites.
X. Cai, H. Xu, S. Xu, Y. Zhang, et al
2022
Earlier work this paper cites.
H.-y. Lu, C. Fan, J. Yang, C. Hu, W. Fang, and X.-j. Wu, “Where to attack: A dynamic locator model for backdoor attack in text classifications,” in Proceedings of the 29th International Conference on Computational Linguistics
2022
Earlier work this paper cites.
Y. Chen, F. Qi, H. Gao, Z. Liu, and M. Sun, “Textual backdoor attacks can be more harmful via two simple tricks,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing
2022
Earlier work this paper cites.
E. Bagdasaryan and V. Shmatikov, “Spinning language models: Risks of propaganda-as-a-service and countermeasures,” in 2022 IEEE Symposium on Security and Privacy (SP)
2022
Cited alongside, same era.
L. Gan, J. Li, T. Zhang, X. Li, Y. Meng, F. Wu, Y. Yang, S. Guo, and C. Fan, “Triggerless backdoor attack for nlp tasks with clean labels,” in Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
2022
Cited alongside, same era.
X. Pan, M. Zhang, B. Sheng, J. Zhu, and M. Yang, “Hidden trigger backdoor attack on { \{ NLP } \} models via linguistic style manipulation,” in 31st USENIX Security Symposium (USENIX Security 22)
2022
Cited alongside, same era.
K. Shao, Y. Zhang, J. Yang, X. Li, and H. Liu, “The triggers that open the nlp model backdoors are hidden in the adversarial samples,” Computers & Security
2022
Cited alongside, same era.
2023
Closest in time.
2023
Closest in time.
S. Zhai, Q. Shen, X. Chen, W. Wang, C. Li, Y. Fang, and Z. Wu, “Ncl: Textual backdoor defense using noise-augmented contrastive learning,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
2023
Closest in time.
P. Li, P. Cheng, F. Li, W. Du, H. Zhao, and G. Liu, “Plmmark: a secure and robust black-box watermarking framework for pre-trained language models,” in Proceedings of the AAAI Conference on Artificial Intelligence
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
X. Chen, Y. Dong, Z. Sun, S. Zhai, Q. Shen, and Z. Wu, “Kallima: A clean-label framework for textual backdoor attacks,” in Computer Security–ESORICS 2022: 27th European Symposium on Research in Computer Security, Copenhagen, Denmark, September 26–30, 2022, Proceedings, Part I
2022
Cited alongside, same era.
L. Jin, Z. Wang, and J. Shang, “Wedef: Weakly supervised backdoor defense for text classification,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing
2022
Cited alongside, same era.
S. Chen, W. Yang, Z. Zhang, X. Bi, and X. Sun, “Expose backdoors on the way: A feature-based efficient defense against textual backdoor attacks,” in Findings of the Association for Computational Linguistics: EMNLP 2022
2022
Cited alongside, same era.
Z. Zhang, L. Lyu, X. Ma, C. Wang, and X. Sun, “Fine-mixing: Mitigating backdoors in fine-tuned language models,” in Findings of the Association for Computational Linguistics: EMNLP 2022
2022
Cited alongside, same era.
B. Zhu, Y. Qin, G. Cui, Y. Chen, W. Zhao, C. Fu, Y. Deng, Z. Liu, J. Wang, W. Wu, et al
2022
Cited alongside, same era.
G. Shen, Y. Liu, G. Tao, Q. Xu, Z. Zhang, S. An, S. Ma, and X. Zhang, “Constrained optimization with dynamic bound-scaling for effective nlp backdoor defense,” in International Conference on Machine Learning
2022
Cited alongside, same era.
W. Lyu, S. Zheng, T. Ma, and C. Chen, “A study of the attention abnormality in trojaned berts,” in Annual Conference of the North American Chapter of the Association for Computational Linguistics
2022
Cited alongside, same era.
T. Yang, H. Wu, B. Yi, G. Feng, and X. Zhang, “Semantic-preserving linguistic steganography by pivot translation and semantic-aware bins coding,” IEEE Transactions on Dependable and Secure Computing
2023
Closest in time.
2024
Closest in time.
2024
Closest in time.
J. Xue, M. Zheng, T. Hua, Y. Shen, Y. Liu, L. Bölöni, and Q. Lou, “Trojllm: A black-box trojan prompt attack on large language models,” Advances in Neural Information Processing Systems
2024
Closest in time.
2024
Closest in time.
W. Du, T. Ju, G. Ren, G. Li, and G. Liu, “Backdoor nlp models via ai-generated text,” in Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)
2024
Closest in time.
W. Du, T. Yuan, H. Zhao, and G. Liu, “Nws: Natural textual backdoor attacks via word substitution,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
X. Li, X. Lu, and P. Li, “Leverage nlp models against other nlp models: Two invisible feature space backdoor attacks,” IEEE Transactions on Reliability
2024
Closest in time.
S. Zhao, L. A. Tuan, J. Fu, J. Wen, and W. Luo, “Exploring clean label backdoor attacks and defense in language models,” IEEE/ACM Transactions on Audio, Speech, and Language Processing
2024
Closest in time.
Y. Li, T. Li, K. Chen, J. Zhang, S. Liu, W. Wang, T. Zhang, and Y. Liu, “Badedit: Backdooring large language models by model editing,” in The Twelfth International Conference on Learning Representations
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
L. Yan, Z. Zhang, G. Tao, K. Zhang, X. Chen, G. Shen, and X. Zhang, “Parafuzz: An interpretability-driven technique for detecting poisoned samples in nlp,” Advances in Neural Information Processing Systems
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Xi, T. Du, C. Li, R. Pang, S. Ji, J. Chen, F. Ma, and T. Wang, “Defending pre-trained language models as few-shot learners against backdoor attacks,” Advances in Neural Information Processing Systems
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Yang, B. Xu, J. M. Zhang, H. J. Kang, J. Shi, J. He, and D. Lo, “Stealthy backdoor attack for code models,” IEEE Transactions on Software Engineering
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
A. Liu, Y. Zhou, X. Liu, T. Zhang, S. Liang, J. Wang, Y. Pu, T. Li, J. Zhang, W. Zhou, et al
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
L. Zhu, R. Ning, J. Li, C. Xin, and H. Wu, “Seer: Backdoor detection for vision-language models through searching target text and image trigger jointly,” in Proceedings of the AAAI Conference on Artificial Intelligence
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
J. Rando and F. Tramèr, “Universal jailbreak backdoors from poisoned human feedback,” in The Twelfth International Conference on Learning Representations
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.