Fetching the paper…
Reading the bibliography…
Natural language processing (NLP) systems have been proven to be vulnerable to backdoor attacks, whereby hidden features (backdoors) are trained into a language model and may only be activated by specific inputs (called triggers), to trick the model into producing unexpected behaviors.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Speech & language processing
Dan Jurafsky. 2000 · 2000
Earlier work this paper cites.
Foundations of Statistical Natural Language Processing
Christopher D. Manning and Hinrich Schütze. 2001 · 2001
Earlier work this paper cites.
BLEU: a Method for Automatic Evaluation of Machine Translation. In Proc. of ACL
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Jauvin. 2003 · 2003
Earlier work this paper cites.
Cutting through the Confusion: A Measurement Study of Homograph Attacks.. In USENIX Annual Technical Conference, General Track . 261–266
Tobias Holgers, David E Watson, and Steven D Gribble. 2006 · 2006
Earlier work this paper cites.
Disinformation on the Web: Impact, Characteristics, and Detection of Wikipedia Hoaxes. In Proc. of WWW
Srijan Kumar, Robert West, and Jure Leskovec. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100, 000+ Questions for Machine Comprehension of Text. In Proc. of EMNLP
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Neural Machine Translation of Rare Words with Subword Units. In Proc. of ACL
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Trojaning Attack on Neural Networks. In Proc. of NDSS
Yingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee, Juan Zhai, Weihang Wang, and Xiangyu Zhang. 2017 · 2017
Earlier work this paper cites.
Universal Adversarial Perturbations. In Proc. of IEEE CVPR
Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, Omar Fawzi, and Pascal Frossard. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need. In Proc. of NeurIPS
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
LEMNA: Explaining Deep Learning based Security Applications. In Proc. of CCS
Wenbo Guo, Dongliang Mu, Jun Xu, Purui Su, Gang Wang, and Xinyu Xing. 2018 · 2018
Earlier work this paper cites.
Manipulating Machine Learning: Poisoning Attacks and Countermeasures for Regression Learning. In Proc. of IEEE S&P
Matthew Jagielski, Alina Oprea, Battista Biggio, Chang Liu, Cristina Nita-Rotaru, and Bo Li. 2018 · 2018
Earlier work this paper cites.
SoK: Security and Privacy in Machine Learning. In Proc. of IEEE EuroS&P
Nicolas Papernot, Patrick D. McDaniel, Arunesh Sinha, and Michael P. Wellman. 2018 · 2018
Earlier work this paper cites.
A Call for Clarity in Reporting BLEU Scores. In Proc. of the Third Conference on Machine Translation: Research Papers
Matt Post. 2018 · 2018
Earlier work this paper cites.
Know What You Don’t Know: Unanswerable Questions for SQuAD. In Proc. of ACL
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
Asking for a Friend: Evaluating Response Biases in Security User Studies. In Proc. of CCS
Elissa M Redmiles, Ziyun Zhu, Sean Kross, Dhruv Kuchhal, Tudor Dumitras, and Michelle L Mazurek. 2018 · 2018
Earlier work this paper cites.
Fast and Effective Robustness Certification. In Proc. of NeurIPS
Gagandeep Singh, Timon Gehr, Matthew Mirman, Markus Püschel, and Martin T. Vechev. 2018 · 2018
Earlier work this paper cites.
Detecting Homoglyph Attacks with a Siamese Neural Network. In Proc. of IEEE Security and Privacy Workshops (SPW)
J. Woodbridge, H. S. Anderson, A. Ahuja, and D. Grant. 2018 · 2018
Earlier work this paper cites.
A Backdoor Attack Against LSTM-Based Text Classification Systems
Jiazhu Dai, Chuanshuai Chen, and Yufeng Li. 2019 · 2019
Earlier work this paper cites.
Why Do Adversarial Attacks Transfer? Explaining Transferability of Evasion and Poisoning Attacks. In Proc. of USENIX Security
Ambra Demontis, Marco Melis, Maura Pintor, Matthew Jagielski, Battista Biggio, Alina Oprea, Cristina Nita-Rotaru, and Fabio Roli. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proc. of NAACL-HLT
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
STRIP: A Defence against Trojan Attacks on Deep Neural Networks. In Proc. of ACSAC
Yansong Gao, Change Xu, Derui Wang, Shiping Chen, Damith C. Ranasinghe, and Surya Nepal. 2019 · 2019
Earlier work this paper cites.
A healthier Twitter: Progress and more to do
D. Hicks and D. Gasca. 2020 · 2019
Earlier work this paper cites.
TextBugger: Generating Adversarial Text Against Real-world Applications. In Proc. of NDSS
Jinfeng Li, Shouling Ji, Tianyu Du, Bo Li, and Ting Wang. 2019 · 2019
Cited alongside, same era.
Poster: Adversarial Examples for Hate Speech Classifiers. In Proc. of CCS
Rajvardhan Oak. 2019 · 2019
Cited alongside, same era.
fairseq: A Fast, Extensible Toolkit for Sequence Modeling. In Proc. of NAACL-HLT 2019: Demonstrations
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Cited alongside, same era.
Defending Neural Backdoors via Generative Distribution Modeling. In Proc. of NeurIPS
Ximing Qiao, Yukun Yang, and Hai Li. 2019 · 2019
Cited alongside, same era.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Adversarial Preprocessing: Understanding and Preventing Image-Scaling Attacks in Machine Learning. In Proc. of USENIX Security
Erwin Quiring, David Klein, Daniel Arp, Martin Johns, and Konrad Rieck. 2020 · 2020
Later among the works it cites.
TBT: Targeted Neural Network Attack with Bit Trojan. In Proc. of IEEE/CVF CVPR
Adnan Siraj Rakin, Zhezhi He, and Deliang Fan. 2020 · 2020
Later among the works it cites.
Don’t Trigger Me! A Triggerless Backdoor Attack Against Deep Neural Networks
Ahmed Salem, Michael Backes, and Yang Zhang. 2020 · 2020
Later among the works it cites.
Dynamic Backdoor Attacks Against Machine Learning Models
Ahmed Salem, Rui Wen, Michael Backes, Shiqing Ma, and Yang Zhang. 2020 · 2020
Later among the works it cites.
Gotta Catch’Em All: Using Honeypots to Catch Adversarial Attacks on Neural Networks. In Proc. of CCS
Shawn Shan, Emily Wenger, Bolun Wang, Bo Li, Haitao Zheng, and Ben Y. Zhao. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mahmood Sharif, Sruti Bhagavatula, Lujo Bauer, and Michael K. Reiter. 2019 · 2019
Cited alongside, same era.
Universal Adversarial Triggers for Attacking and Analyzing NLP. In Proc. of EMNLP-IJCNLP
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2019 · 2019
Cited alongside, same era.
Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural Networks. In Proc. IEEE S&P
Bolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li, Bimal Viswanath, Haitao Zheng, and Ben Y. Zhao. 2019 · 2019
Cited alongside, same era.
Rendered Private: Making GLSL Execution Uniform to Prevent WebGL-based Browser Fingerprinting. In Proc. of USENIX Security
Shujiang Wu, Song Li, Yinzhi Cao, and Ningfei Wang. 2019 · 2019
Cited alongside, same era.
Analyzing Information Leakage of Updates to Natural Language Models. In Proc. of CCS
Santiago Zanella Béguelin, Lukas Wutschitz, and Shruti Tople et al. 2020 · 2020
Cited alongside, same era.
Extracting Training Data from Large Language Models
Nicholas Carlini, Florian Tramer, and Eric Wallace et al. 2020 · 2020
Cited alongside, same era.
BadNL: Backdoor Attacks Against NLP Models
Xiaoyi Chen, Ahmed Salem, Michael Backes, Shiqing Ma, and Yang Zhang. 2020 · 2020
Cited alongside, same era.
Adversarial Semantic Collisions. In Proc. of EMNLP
Congzheng Song, Alexander M. Rush, and Vitaly Shmatikov. 2020 · 2020
Later among the works it cites.
Bypassing Backdoor Detection Algorithms in Deep Learning. In Proc. of IEEE EuroS&P
Te Juin Lester Tan and Reza Shokri. 2020 · 2020
Later among the works it cites.
Rapidly Bootstrapping a Question Answering Dataset for COVID-19
Raphael Tang, Rodrigo Nogueira, Edwin Zhang, Nikhil Gupta, Phuong Cam, Kyunghyun Cho, and Jimmy Lin. 2020 · 2020
Later among the works it cites.
Imitation Attacks and Defenses for Black-box Machine Translation Systems. In Proc. of EMNLP
Eric Wallace, Mitchell Stern, and Dawn Song. 2020 · 2020
Later among the works it cites.
Detecting AI Trojans Using Meta Neural Analysis. In Proc. of IEEE S&P
Xiaojun Xu, Qi Wang, Huichen Li, Nikita Borisov, Carl A. Gunter, and Bo Li. 2020 · 2020
Later among the works it cites.
Backdoor Attacks to Graph Neural Networks
Zaixi Zhang, Jinyuan Jia, Binghui Wang, and Neil Zhenqiang Gong. 2020 · 2020
Later among the works it cites.
Blind Backdoors in Deep Learning Models. In Proc. of USENIX Security
Eugene Bagdasaryan and Vitaly Shmatikov. 2021 · 2021
Closest in time.
Data Poisoning Attacks to Local Differential Privacy Protocols. In Proc. of USENIX Security
Xiaoyu Cao, Jinyuan Jia, and Neil Zhenqiang Gong. 2021 · 2021
Closest in time.
Deep Feature Space Trojan Attack of Neural Networks by Controlled Detoxification. In Proc. of AAAI
Siyuan Cheng, Yingqi Liu, Shiqing Ma, and Xiangyu Zhang. 2021 · 2021
Closest in time.
Confusables
Unicode Consortium. 2020 · 2021
Closest in time.
Data Poisoning Attacks to Deep Learning Based Recommender Systems. In Proc. of NDSS
Hai Huang, Jiaming Mu, Neil Zhenqiang Gong, Qi Li, Bin Liu, and Mingwei Xu. 2021 · 2021
Closest in time.
Intrinsic Certified Robustness of Bagging against Data Poisoning Attacks. In Proc. of AAAI
Jinyuan Jia, Xiaoyu Cao, and Neil Zhenqiang Gong. 2021 · 2021
Closest in time.
The Audio Auditor: User-Level Membership Inference in Internet of Things Voice Services
Yuantian Miao, Minhui Xue, Chao Chen, Lei Pan, Jun Zhang, Benjamin Zi Hao Zhao, Dali Kaafar, and Yang Xiang. 2021 · 2021
Closest in time.
WaNet - Imperceptible Warping-based Backdoor Attack
Anh Nguyen and Anh Tran. 2021 · 2021
Closest in time.
Infobert: Improving robustness of language models from an information theoretic perspective. In Proc. of ICLR
Boxin Wang, Shuohang Wang, Yu Cheng, Zhe Gan, Ruoxi Jia, Bo Li, and Jingjing Liu. 2021 · 2021
Closest in time.
With Great Dispersion Comes Greater Resilience: Efficient Poisoning Attacks and Defenses for Linear Regression Models
Jialin Wen, Benjamin Zi Hao Zhao, Minhui Xue, Alina Oprea, and Haifeng Qian. 2021 · 2021
Closest in time.
Graph Backdoor. In Proc. of USENIX Security
Zhaohan Xi, Ren Pang, Shouling Ji, and Ting Wang. 2021 · 2021
Closest in time.
Targeted Poisoning Attacks on Black-Box Neural Machine Translation. In Proc. of WWW
Chang Xu, Jun Wang, Yuqing Tang, Francisco Guzman, Benjamin IP Rubinstein, and Trevor Cohn. 2021 · 2021
Closest in time.
Trojaning Language Models for Fun and Profit. In Proc. of IEEE EuroS&P
Xinyang Zhang, Zheng Zhang, and Ting Wang. 2021 · 2021
Closest in time.