Fetching the paper…
Reading the bibliography…
Backdoor Attacks have been a serious vulnerability against Large Language Models (LLMs).
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R.; Perelygin, A.; Wu, J.; Chuang, J.; Manning, C. D.; Ng, A. Y.; and Potts, C. 2013 · 2013
Earlier work this paper cites.
Rusu, A. A.; Colmenarejo, S. G.; Gulcehre, C.; Desjardins, G.; Kirkpatrick, J.; Pascanu, R.; Mnih, V.; Kavukcuoglu, K.; and Hadsell, R. 2015 · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Zhang, X.; Zhao, J.; and LeCun, Y. 2015 · 2015
Earlier work this paper cites.
A backdoor attack against lstm-based text classification systems
Dai, J.; Chen, C.; and Li, Y. 2019 · 2019
Earlier work this paper cites.
Strip: A defence against trojan attacks on deep neural networks
Gao, Y.; Xu, C.; Wang, D.; Chen, S.; Ranasinghe, D. C.; and Nepal, S. 2019 · 2019
Earlier work this paper cites.
Universal Adversarial Triggers for Attacking and Analyzing NLP
Wallace, E.; Feng, S.; Kandpal, N.; Gardner, M.; and Singh, S. 2019 · 2019
Earlier work this paper cites.
TinyBERT: Distilling BERT for Natural Language Understanding
Jiao, X.; Yin, Y.; Shang, L.; Jiang, X.; Chen, X.; Li, L.; Wang, F.; and Liu, Q. 2020 · 2020
Earlier work this paper cites.
Weight Poisoning Attacks on Pretrained Models
Kurita, K.; Michel, P.; and Neubig, G. 2020 · 2020
Earlier work this paper cites.
Inoculating against fake news about COVID-19
van Der Linden, S.; Roozenbeek, J.; and Compton, J. 2020 · 2020
Earlier work this paper cites.
SemEval-2020 Task 12: Multilingual Offensive Language Identification in Social Media (OffensEval 2020)
Zampieri, M.; Nakov, P.; Rosenthal, S.; Atanasova, P.; Karadzhov, G.; Mubarak, H.; Derczynski, L.; Pitenis, Z.; and Çöltekin, Ç. 2020 · 2020
Earlier work this paper cites.
T-miner: A generative approach to defend against trojan attacks on dnn-based text classification
Azizi, A.; Tahmid, I. A.; Waheed, A.; Mangaokar, N.; Pu, J.; Javed, M.; Reddy, C. K.; and Viswanath, B. 2021 · 2021
Earlier work this paper cites.
GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow
Black, S.; Leo, G.; Wang, P.; Leahy, C.; and Biderman, S. 2021 · 2021
Earlier work this paper cites.
Anti-distillation backdoor attacks: Backdoors can really survive in knowledge distillation
Ge, Y.; Wang, Q.; Zheng, B.; Zhuang, X.; Li, Q.; Shen, C.; and Wang, C. 2021 · 2021
Earlier work this paper cites.
Knowledge distillation: A survey
Gou, J.; Yu, B.; Maybank, S. J.; and Tao, D. 2021 · 2021
Earlier work this paper cites.
ONION: A Simple and Effective Defense Against Textual Backdoor Attacks
Qi, F.; Chen, Y.; Li, M.; Yao, Y.; Liu, Z.; and Sun, M. 2021a · 2021
Earlier work this paper cites.
Mind the Style of Text! Adversarial and Backdoor Attacks Based on Text Style Transfer
Qi, F.; Chen, Y.; Zhang, X.; Li, M.; Liu, Z.; and Sun, M. 2021b · 2021
Earlier work this paper cites.
Backdoor Pre-trained Models Can Transfer to All
Shen, L.; Ji, S.; Zhang, X.; Li, J.; Chen, J.; Shi, J.; Fang, C.; Yin, J.; and Wang, T. 2021 · 2021
Earlier work this paper cites.
MiniLMv2: Multi-Head Self-Attention Relation Distillation for Compressing Pretrained Transformers
Wang, W.; Bao, H.; Huang, S.; Dong, L.; and Wei, F. 2021 · 2021
Earlier work this paper cites.
Dual Contrastive Learning: Text Classification via Label-Aware Data Augmentation
Chen, Q.; Zhang, R.; Zheng, Y.; and Mao, Y. 2022 · 2022
Cited alongside, same era.
Multitask Prompted Training Enables Zero-Shot Task Generalization
Sanh, V.; Webson, A.; Raffel, C.; Bach, S. H.; Sutawika, L.; Alyafeai, Z.; Chaffin, A.; Stiegler, A.; Le Scao, T.; Raja, A.; et al. 2022 · 2022
Cited alongside, same era.
A comprehensive survey on poisoning attacks and countermeasures in machine learning
Tian, Z.; Cui, L.; Liang, J.; and Yu, S. 2022 · 2022
Cited alongside, same era.
Finetuned Language Models are Zero-Shot Learners
Wei, J.; Bosma, M.; Zhao, V.; Guu, K.; Yu, A. W.; Lester, B.; Du, N.; Dai, A. M.; and Le, Q. V. 2022 · 2022
Cited alongside, same era.
OPT: Open Pre-trained Transformer Language Models
Zhang, S.; Roller, S.; Goyal, N.; Artetxe, M.; Chen, M.; Chen, S.; Dewan, C.; Diab, M.; Li, X.; Lin, X. V.; Mihaylov, T.; Ott, M.; Shleifer, S.; Shuster, K.; Simig, D.; Koura, P. S.; Sridhar, A.; Wang, T.; and Zettlemoyer, L. 2022 · 2022
Stanford alpaca: an instruction-following llama model (2023)
Taori, R.; Gulrajani, I.; Zhang, T.; Dubois, Y.; Li, X.; Guestrin, C.; Liang, P.; and Hashimoto, T. B. 2023 · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Team, G.; Anil, R.; Borgeaud, S.; Wu, Y.; Alayrac, J.-B.; Yu, J.; Soricut, R.; Schalkwyk, J.; Dai, A. M.; Hauth, A.; et al. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Later among the works it cites.
BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models
Xiang, Z.; Jiang, F.; Xiong, Z.; Ramasubramanian, B.; Poovendran, R.; and Li, B. 2023 · 2023
Later among the works it cites.
Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models
Zhao, S.; Wen, J.; Luu, A.; Zhao, J.; and Fu, J. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Cited alongside, same era.
Anil, R.; Dai, A. M.; Firat, O.; Johnson, M.; Lepikhin, D.; Passos, A.; Shakeri, S.; Taropa, E.; Bailey, P.; Chen, Z.; et al. 2023 · 2023
Cited alongside, same era.
Pythia: A suite for analyzing large language models across training and scaling
Biderman, S.; Schoelkopf, H.; Anthony, Q. G.; Bradley, H.; O’Brien, K.; Hallahan, E.; Khan, M. A.; Purohit, S.; Prashanth, U. S.; Raff, E.; et al. 2023 · 2023
Cited alongside, same era.
Cheng, P.; Wu, Z.; Du, W.; and Liu, G. 2023 · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; et al. 2023 · 2023
Cited alongside, same era.
Unleashing cheapfakes through trojan plugins of large language models
Dong, T.; Chen, G.; Li, S.; Xue, M.; Holland, R.; Meng, Y.; Liu, Z.; and Zhu, H. 2023 · 2023
Cited alongside, same era.
Uor: Universal backdoor attacks on pre-trained language models
Du, W.; Li, P.; Li, B.; Zhao, H.; and Liu, G. 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
A survey on model compression for large language models
Zhu, X.; Li, J.; Liu, Y.; Ma, C.; and Wang, W. 2023 · 2023
Later among the works it cites.
Mitigating Backdoor Attacks in Pre-trained Encoders via Self-supervised Knowledge Distillation
Bie, R.; Jiang, J.; Xie, H.; Guo, Y.; Miao, Y.; and Jia, X. 2024 · 2024
Closest in time.
Anti-Backdoor Model: A Novel Algorithm To Remove Backdoors in a Non-invasive Way
Chen, C.; Hong, H.; Xiang, T.; and Xie, M. 2024 · 2024
Closest in time.
Safe RLHF: Safe Reinforcement Learning from Human Feedback
Dai, J.; Pan, X.; Sun, R.; Ji, J.; Xu, X.; Liu, M.; Wang, Y.; and Yang, Y. 2024 · 2024
Closest in time.
Composite Backdoor Attacks Against Large Language Models
Huang, H.; Zhao, Z.; Backes, M.; Shen, Y.; and Zhang, Y. 2024 · 2024
Closest in time.
BadEdit: Backdooring Large Language Models by Model Editing
Li, Y.; Li, T.; Chen, K.; Zhang, J.; Liu, S.; Wang, W.; Zhang, T.; and Liu, Y. 2024 · 2024
Closest in time.
Backdoor attacks on dense passage retrievers for disseminating misinformation
Long, Q.; Deng, Y.; Gan, L.; Wang, W.; and Pan, S. J. 2024 · 2024
Closest in time.
Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint
Xiong, W.; Dong, H.; Ye, C.; Wang, Z.; Zhong, H.; Ji, H.; Jiang, N.; and Zhang, T. 2024 · 2024
Closest in time.
From Toxic to Trustworthy: Using Self-Distillation and Semi-supervised Methods to Refine Neural Networks
Zhang, X.; Zheng, B.; Hu, J.; Li, C.; and Bai, X. 2024 · 2024
Closest in time.
Exploring Clean Label Backdoor Attacks and Defense in Language Models
Zhao, S.; Tuan, L. A.; Fu, J.; Wen, J.; and Luo, W. 2024 · 2024
Closest in time.
Piccolo: Exposing complex backdoors in nlp transformer models
Liu, Y.; Shen, G.; Tao, G.; An, S.; Ma, S.; and Zhang, X. 2022 · 2042
Closest in time.
Backdoor NLP Models via AI-Generated Text
Du, W.; Ju, T.; Ren, G.; Li, G.; and Liu, G. 2024 · 2079
Closest in time.