Fetching the paper…
Reading the bibliography…
Despite being widely applied due to their exceptional capabilities, Large Language Models (LLMs) have been proven to be vulnerable to backdoor attacks.
Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning
Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin A Raffel. 2022 · 1965
Earlier work this paper cites.
The information bottleneck method
Naftali Tishby, Fernando C Pereira, and William Bialek. 2000 · 2000
Earlier work this paper cites.
Mining and summarizing customer reviews
Minqing Hu and Bing Liu. 2004 · 2004
Earlier work this paper cites.
A tutorial on the cross-entropy method
Pieter-Tjerk De Boer, Dirk P Kroese, Shie Mannor, and Reuven Y Rubinstein. 2005 · 2005
Earlier work this paper cites.
Ape210k: A large-scale and template-rich dataset of math word problems
Wei Zhao, Mingyue Shang, Yang Liu, Liang Wang, and Jingming Liu. 2020 · 2009
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Deep learning and the information bottleneck principle
Naftali Tishby and Noga Zaslavsky. 2015 · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
Badnets: Identifying vulnerabilities in the machine learning model supply chain
Tianyu Gu, Brendan Dolan-Gavitt, and Siddharth Garg. 2017 · 2017
Earlier work this paper cites.
A backdoor attack against lstm-based text classification systems
Jiazhu Dai, Chuanshuai Chen, and Yufeng Li. 2019 · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Earlier work this paper cites.
Knowledge distillation for multi-task learning
Wei-Hong Li and Hakan Bilen. 2020 · 2020
Earlier work this paper cites.
Mitigating backdoor attacks in lstm-based text classification systems by backdoor keyword identification
Chuanshuai Chen and Jiazhu Dai. 2021 · 2021
Earlier work this paper cites.
Deep feature space trojan attack of neural networks by controlled detoxification
Siyuan Cheng, Yingqi Liu, Shiqing Ma, and Xiangyu Zhang. 2021 · 2021
Earlier work this paper cites.
Anti-distillation backdoor attacks: Backdoors can really survive in knowledge distillation
Yunjie Ge, Qian Wang, Baolin Zheng, Xinlu Zhuang, Qi Li, Chao Shen, and Cong Wang. 2021 · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. 2021 · 2021
Earlier work this paper cites.
Fedpara: Low-rank hadamard product for communication-efficient federated learning
Nam Hyeon-Woo, Moon Ye-Bin, and Tae-Hyun Oh. 2021 · 2021
Earlier work this paper cites.
Comparing kullback-leibler divergence and mean squared error loss in knowledge distillation
Taehyeon Kim, Jaehoon Oh, NakYil Kim, Sangwook Cho, and Se-Young Yun. 2021 · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021 · 2021
Earlier work this paper cites.
Backdoor attacks on pre-trained models by layerwise weight poisoning
Linyang Li, Demin Song, Xiaonan Li, Jiehang Zeng, Ruotian Ma, and Xipeng Qiu. 2021a · 2021
Earlier work this paper cites.
Hidden backdoors in human-centric language models
Shaofeng Li, Hui Liu, Tian Dong, Benjamin Zi Hao Zhao, Minhui Xue, Haojin Zhu, and Jialiang Lu. 2021b · 2021
Earlier work this paper cites.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang. 2021 · 2021
Earlier work this paper cites.
Onion: A simple and effective defense against textual backdoor attacks
Fanchao Qi, Yangyi Chen, Mukai Li, Yuan Yao, Zhiyuan Liu, and Maosong Sun. 2021a · 2021
Earlier work this paper cites.
Concealed data poisoning attacks on nlp models
Eric Wallace, Tony Zhao, Shi Feng, and Sameer Singh. 2021 · 2021
Cited alongside, same era.
Badprompt: Backdoor attacks on continuous prompts
Xiangrui Cai, Sihan Xu, Ying Zhang, Xiaojie Yuan, et al. 2022 · 2022
Cited alongside, same era.
Kallima: A clean-label framework for textual backdoor attacks
Xiaoyi Chen, Yinpeng Dong, Zeyu Sun, Shengfang Zhai, Qingni Shen, and Zhonghai Wu. 2022 · 2022
Cited alongside, same era.
Triggerless backdoor attack for nlp tasks with clean labels
Leilei Gan, Jiwei Li, Tianwei Zhang, Xiaoya Li, Yuxian Meng, Fei Wu, Yi Yang, Shangwei Guo, and Chun Fan. 2022 · 2022
Cited alongside, same era.
Badhash: Invisible backdoor attacks against deep hashing with clean label
Shengshan Hu, Ziqi Zhou, Yechao Zhang, Leo Yu Zhang, Yifeng Zheng, et al. 2022 · 2022
Cited alongside, same era.
Backdoor attack against nlp models with robustness-aware perturbation defense
Prompt as triggers for backdoor attack: Examining the vulnerability in language models
Shuai Zhao, Jinming Wen, Anh Tuan Luu, Junbo Zhao, and Jie Fu. 2023b · 2023
Later among the works it cites.
Adfl: Defending backdoor attacks in federated learning via adversarial distillation
Chengcheng Zhu, Jiale Zhang, Xiaobing Sun, Bing Chen, and Weizhi Meng. 2023 · 2023
Later among the works it cites.
Llama 3 model card
AI@Meta. 2024 · 2024
Closest in time.
Mitigating backdoor attacks in pre-trained encoders via self-supervised knowledge distillation
Rongfang Bie, Jinxiu Jiang, Hongcheng Xie, Yu Guo, Yinbin Miao, and Xiaohua Jia. 2024 · 2024
Closest in time.
Robust knowledge distillation based on feature variance against backdoored teacher model
Jinyin Chen, Xiaoming Zhao, Haibin Zheng, Xiao Li, Sheng Xiang, and Haifeng Guo. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shaik Mohammed Maqsood, Viveros Manuela Ceron, and Addluri GowthamKrishna. 2022 · 2022
Cited alongside, same era.
Improving neural cross-lingual abstractive summarization via employing optimal transport distance for knowledge distillation
Thong Thanh Nguyen and Anh Tuan Luu. 2022 · 2022
Cited alongside, same era.
A knowledge distillation-based backdoor attack in federated learning
Yifan Wang, Wei Fan, Keke Yang, Naji Alhusaini, and Jing Li. 2022 · 2022
Cited alongside, same era.
Exploring the universal vulnerability of prompt-based learning paradigm
Lei Xu, Yangyi Chen, Ganqu Cui, Hongcheng Gao, and Zhiyuan Liu. 2022 · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022 · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Cited alongside, same era.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, et al. 2023 · 2023
Cited alongside, same era.
Pengzhou Cheng, Zongru Wu, Tianjie Ju, Wei Du, and Zhuosheng Zhang Gongshen Liu. 2024 · 2024
Closest in time.
Comprehensive assessment of jailbreak attacks against llms
Junjie Chu, Yugeng Liu, Ziqing Yang, Xinyue Shen, Michael Backes, and Yang Zhang. 2024 · 2024
Closest in time.
Light-peft: Lightening parameter-efficient fine-tuning via early pruning
Naibin Gu, Peng Fu, Xiyu Liu, Bowen Shen, Zheng Lin, and Weiping Wang. 2024 · 2024
Closest in time.
Artwork protection against neural style transfer using locally adaptive adversarial color attack
Zhongliang Guo, Kaixuan Wang, Weiye Li, Yifei Qian, Ognjen Arandjelović, and Lei Fang. 2024 · 2024
Closest in time.
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024 · 2024
Closest in time.
Chatgpt as an attack tool: Stealthy textual backdoor attack via blackbox generative model trigger
Jiazhao Li, Yijin Yang, Zhuofeng Wu, VG Vinod Vydiswaran, and Chaowei Xiao. 2024a · 2024
Closest in time.
Backdoor attacks on dense passage retrievers for disseminating misinformation
Quanyu Long, Yue Deng, LeiLei Gan, Wenya Wang, and Sinno Jialin Pan. 2024 · 2024
Closest in time.
On the affinity, rationality, and diversity of hierarchical topic modeling
Xiaobao Wu, Fengjun Pan, Thong Nguyen, Yichao Feng, Chaoqun Liu, Cong-Duy Nguyen, and Anh Tuan Luu. 2024 · 2024
Closest in time.
Atlantis: Aesthetic-oriented multiple granularities fusion network for joint multimodal aspect-based sentiment analysis
Luwei Xiao, Xingjiao Wu, Junjie Xu, Weijie Li, Cheng Jin, and Liang He. 2024 · 2024
Closest in time.
Trojllm: A black-box trojan prompt attack on large language models
Jiaqi Xue, Mengxin Zheng, Ting Hua, Yilin Shen, Yepeng Liu, Ladislau Bölöni, and Qian Lou. 2024 · 2024
Closest in time.
Defending against weight-poisoning backdoor attacks for parameter-efficient fine-tuning
Shuai Zhao, Leilei Gan, Luu Anh Tuan, Jie Fu, Lingjuan Lyu, Meihuizi Jia, and Jinming Wen. 2024a · 2024
Closest in time.
Universal vulnerabilities in large language models: Backdoor attacks for in-context learning
Shuai Zhao, Meihuizi Jia, Luu Anh Tuan, Fengjun Pan, and Jinming Wen. 2024b · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, et al. 2024 · 2024
Closest in time.
Weak-to-strong search: Align large language models via searching over small language models
Zhanhui Zhou, Zhixuan Liu, Jie Liu, Zhichen Dong, Chao Yang, and Yu Qiao. 2024 · 2024
Closest in time.
Jinming Wen, Xinyi Wu, Shuai Zhao, Yanhao Jia, and Yuwen Li. 2025 · 2025
Closest in time.
Exploring cognitive and aesthetic causality for multimodal aspect-based sentiment analysis
Luwei Xiao, Rui Mao, Shuai Zhao, Qika Lin, Yanhao Jia, Liang He, and Erik Cambria. 2025 · 2025
Closest in time.
Unlearning backdoor attacks for llms with weak-to-strong knowledge distillation
Shuai Zhao, Xiaobao Wu, Cong-Duy Nguyen, Yanhao Jia, Meihuizi Jia, Yichao Feng, and Luu Anh Tuan. 2025b · 2025
Closest in time.
Can adversarial weight perturbations inject neural backdoors
Siddhant Garg, Adarsh Kumar, Vibhor Goel, and Yingyu Liang. 2020 · 2032
Closest in time.