Fetching the paper…
Reading the bibliography…
Backdoors can be injected into NLP models to induce misbehavior when the input text contains a specific feature, known as a trigger, which the attacker secretly selects.
Foundations of statistical natural language processing
Christopher D. Manning and Hinrich Schütze · 2001
Earlier work this paper cites.
Concentration inequalities: A nonasymptotic theory of independence
Stephane Boucheron, Gabor Lugosi, and Pascal Massart · 2013
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2015
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Adversarial example generation with syntactically controlled paraphrase networks
Mohit Iyyer, John Wieting, Kevin Gimpel, and Luke Zettlemoyer · 2018
Earlier work this paper cites.
Model-reuse attacks on deep learning systems
Yujie Ji, Xinyang Zhang, Shouling Ji, Xiapu Luo, and Ting Wang · 2018
Earlier work this paper cites.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2018
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2018
Earlier work this paper cites.
A backdoor attack against lstm-based text classification system
Jiazhu Dai, Chuanshuai Chen, and Yufeng Li · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
ABS: Scanning neural networks for back-doors by artificial brain stimulation
Yingqi Liu, Wen-Chuan Lee, Guanhong Tao, Shiqing Ma, Yousra Aafer, and Xiangyu Zhang · 2019
Earlier work this paper cites.
RoBERTa: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2019
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2019
Earlier work this paper cites.
Regula sub-rosa: Latent backdoor attacks on deep neural networks
Yuanshun Yao, Huiying Li, Haitao Zheng, and Ben Y. Zhao · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, and Nick Ryder et al · 2020
Earlier work this paper cites.
Plug and play language models: A simple approach to controlled text generation
Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu · 2020
Earlier work this paper cites.
Fantastic generalization measures and where to find them
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 2020
Earlier work this paper cites.
Reformulating unsupervised style transfer as paraphrase generation
Kalpesh Krishna, John Wieting, and Mohit Iyyer · 2020
Cited alongside, same era.
Weight poisoning attacks on pre-trained models
Keita Kurita, Paul Michel, and Graham Neubig · 2020
Cited alongside, same era.
CC-News-En: A large english news corpus
Joel Mackenzie, Rodger Benham, Matthias Petri, Johanne R. Trippas, J. Shane Culpepper, and Alistair Moffat · 2020
Cited alongside, same era.
ONION: A simple and effective defense against textual backdoor attacks
Fanchao Qi, Yangyi Chen, Mukai Li, Yuan Yao, and Maosong Sun Zhiyuan Liu · 2020
Cited alongside, same era.
T-Miner: A generative approach to defend against trojan attacks on dnn-based text classification
Ahmadreza Azizi, Ibrahim Asadullah Tahmid, Asim Waheed, Neal Mangaokar, Jiameng Pu, Mobin Javed, Chandan K. Reddy, and Bimal Viswanath · 2021
Cited alongside, same era.
Blind backdoors in deep learning models
Spinning language models: Risks of propaganda-as-a-service and countermeasures
Eugene Bagdasaryan and Vitaly Shmatikov · 2022
Later among the works it cites.
LoRA: Low-rank adaptation of large language models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Later among the works it cites.
PICCOLO: Exposing complex backdoors in nlp transformer models
Yingqi Liu, Guangyu Shen, Guanhong Tao, Shengwei An, Shiqing Ma, and Xiangyu Zhang · 2022
Later among the works it cites.
Hidden trigger backdoor attack on nlp models via linguistic style manipulation
Xudong Pan, Mi Zhang, Beina Sheng, Jiaming Zhu, and Min Yang · 2022
Later among the works it cites.
Constrained optimization with dynamic bound-scaling for effective nlp backdoor defense
Guangyu Shen, Yingqi Liu, Guanhong Tao, Qiuling Xu, Zhuo Zhang, Shengwei An, Shiqing Ma, and Xiangyu Zhang · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Eugene Bagdasaryan and Vitaly Shmatikov · 2021
Cited alongside, same era.
GPT-Neo: Large scale autoregressive language modeling with mesh-tensorflow, 2021
Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman · 2021
Cited alongside, same era.
Badnl: Backdoor attacks against nlp models with semantic-preserving improvements
Xiaoyi Chen, Ahmed Salem, Dingfan Chen, Michael Backes, Shiqing Ma, Qingni Shen, Zhonghai Wu, and Yang Zhang · 2021
Cited alongside, same era.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2021
Cited alongside, same era.
Controlled text generation as continuous optimization with multiple constraints
Sachin Kumar, Eric Malmi, Aliaksei Severyn, and Yulia Tsvetkov · 2021
Cited alongside, same era.
Hidden backdoors in human-centric language models
Shaofeng Li, Hui Liu, Tian Dong, Benjamin Zi Hao Zhao, Minhui Xue, Haojin Zhu, and Jialiang Lu · 2021
Cited alongside, same era.
BFClass: A backdoor-free text classification framework
Zichao Li, Dheeraj Mekala, Chengyu Dong, and Jingbo Shang · 2021
Cited alongside, same era.
Training with more confidence: Mitigating injected and natural backdoors during training
Zhenting Wang, Hailun Ding, Juan Zhai, and Shiqing Ma · 2022
Later among the works it cites.
Hero: Hessian-enhanced robust optimization for unifying and improving generalization and quantization performance
Huanrui Yang, Xiaoxuan Yang, Neil Zhenqiang Gong, and Yiran Chen · 2022
Later among the works it cites.
OPT: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, and Naman Goyal et al · 2022
Later among the works it cites.
https://crfm.stanford.edu/2023/03/13/alpaca.html
Alpaca · 2023
Later among the works it cites.
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, and Quentin Anthony et al · 2023
Later among the works it cites.
The philosopher’s stone: Trojaning plugins of large language models
Tian Dong, Minhui Xue, Guoxing Chen, Rayne Holland, Shaofeng Li, Yan Meng, Zhen Liu, and Haojin Zhu · 2023
Later among the works it cites.
FreeEagle: Detecting complex neural trojans in data-free cases
Chong Fu, Xuhong Zhang, Shouling Ji, Ting Wang, Peng Lin, Yanghe Feng, and Jianwei Yin · 2023
Later among the works it cites.
Trojtext: Test-time invisible textual trojan insertion
Yepeng Liu, Bo Feng, and Qian Lou · 2023
Later among the works it cites.
The “Beatrix” resurrections: Robust backdoor detection via gram matrices
Wanlun Ma, Derui Wang, Ruoxi Sun, Minhui Xue, Sheng Wen, and Yang Xiang · 2023
Later among the works it cites.
How does sharpness-aware minimization minimize sharpness?
Kaiyue Wen, Tengyu Ma, and Zhiyuan Li · 2023
Later among the works it cites.
ParaFuzz: An interpretability-driven technique for detecting poisoned samples in nlp
Lu Yan, Zhuo Zhang, Guanhong Tao, Kaiyuan Zhang, Xuan Chen, Guangyu Shen, and Xiangyu Zhang · 2023
Later among the works it cites.
How to sift out a clean data subset in the presence of data poisoning?
Yi Zeng, Minzhou Pan, Himanshu Jahagirdar, Ming Jin, Lingjuan Lyu, and Ruoxi Jia · 2023
Later among the works it cites.
TrojanPuzzle: Covertly poisoning code-suggestion models
Hojjat Aghakhani, Wei Dai, Andre Manoel, Xavier Fernandes, Anant Kharkar, Christopher Kruegel, Giovanni Vigna, David Evans, Ben Zorn, and Robert Sim · 2024
Closest in time.
Sleeper agents: Training deceptive llms that persist through safety training
Evan Hubinger, Carson Denison, and Jesse Mu et al · 2024
Closest in time.
Mm-bd: Post-training detection of backdoor attacks with arbitrary backdoor pattern types using a maximum margin statistic
Hang Wang, Zhen Xiang, David J. Miller, and George Kesidis · 2024
Closest in time.
BadChain: Backdoor chain-of-thought prompting for large language models
Zhen Xiang, Fengqing Jiang, Zidi Xiong, Bhaskar Ramasubramanian, Radha Poovendran, and Bo Li · 2024
Closest in time.