Fetching the paper…
Reading the bibliography…
Backdoor attacks are commonly executed by contaminating training data, such that a trigger can activate predetermined harmful effects during the test phase.
Bleu: a method for automatic evaluation of machine translation
Papineni, Kishore, Roukos, Salim, Ward, Todd, and Zhu, Wei-Jing · 2002
Earlier work this paper cites.
Defending against backdoor attack on deep neural networks
Xu, Kaidi, Liu, Sijia, Chen, Pin-Yu, Zhao, Pu, and Lin, Xue · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, Chin-Yew · 2004
Earlier work this paper cites.
Evasion attacks against machine learning at test time
Biggio, Battista, Corona, Igino, Maiorca, Davide, Nelson, Blaine, Šrndić, Nedim, Laskov, Pavel, Giacinto, Giorgio, and Roli, Fabio · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, Tsung-Yi, Maire, Michael, Belongie, Serge, Hays, James, Perona, Pietro, Ramanan, Deva, Dollár, Piotr, and Zitnick, C Lawrence · 2014
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, Christian, Zaremba, Wojciech, Sutskever, Ilya, Bruna, Joan, Erhan, Dumitru, Goodfellow, Ian, and Fergus, Rob · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Goodfellow, Ian J, Shlens, Jonathon, and Szegedy, Christian · 2015
Earlier work this paper cites.
Brown, Tom B, Mané, Dandelion, Roy, Aurko, Abadi, Martín, and Gilmer, Justin · 2017
Earlier work this paper cites.
Targeted backdoor attacks on deep learning systems using data poisoning
Chen, Xinyun, Liu, Chang, Li, Bo, Lu, Kimberly, and Song, Dawn · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Yash, Khot, Tejas, Summers-Stay, Douglas, Batra, Dhruv, and Parikh, Devi · 2017
Earlier work this paper cites.
Badnets: Identifying vulnerabilities in the machine learning model supply chain
Gu, Tianyu, Dolan-Gavitt, Brendan, and Garg, Siddharth · 2017
Earlier work this paper cites.
Universal adversarial perturbations against semantic image segmentation
Hendrik Metzen, Jan, Chaithanya Kumar, Mummadi, Brox, Thomas, and Fischer, Volker · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, Ranjay, Zhu, Yuke, Groth, Oliver, Johnson, Justin, Hata, Kenji, Kravitz, Joshua, Chen, Stephanie, Kalantidis, Yannis, Li, Li-Jia, Shamma, David A, et al · 2017
Earlier work this paper cites.
Adversarial examples in the physical world
Kurakin, Alexey, Goodfellow, Ian, and Bengio, Samy · 2017
Earlier work this paper cites.
Universal adversarial perturbations
Moosavi-Dezfooli, Seyed-Mohsen, Fawzi, Alhussein, Fawzi, Omar, and Frossard, Pascal · 2017
Earlier work this paper cites.
Fast feature fool: A data independent approach to universal adversarial perturbations
Mopuri, Konda Reddy, Garg, Utsav, and Babu, R Venkatesh · 2017
Earlier work this paper cites.
Detecting backdoor attacks on deep neural networks by activation clustering
Chen, Bryant, Carvalho, Wilka, Baracaldo, Nathalie, Ludwig, Heiko, Edwards, Benjamin, Lee, Taesung, Molloy, Ian, and Srivastava, Biplav · 2018
Earlier work this paper cites.
Boosting adversarial attacks with momentum
Dong, Yinpeng, Liao, Fangzhou, Pang, Tianyu, Su, Hang, Zhu, Jun, Hu, Xiaolin, and Li, Jianguo · 2018
Earlier work this paper cites.
Robust physical-world attacks on deep learning visual classification
Eykholt, Kevin, Evtimov, Ivan, Fernandes, Earlence, Li, Bo, Rahmati, Amir, Xiao, Chaowei, Prakash, Atul, Kohno, Tadayoshi, and Song, Dawn · 2018
Earlier work this paper cites.
Backdoor embedding in convolutional neural network models via invisible perturbation
Liao, Cong, Zhong, Haoti, Squicciarini, Anna, Zhu, Sencun, and Miller, David · 2018
Earlier work this paper cites.
Dpatch: An adversarial patch attack on object detectors
Liu, Xin, Yang, Huanrui, Liu, Ziwei, Song, Linghao, Li, Hai, and Chen, Yiran · 2018
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Madry, Aleksander, Makelov, Aleksandar, Schmidt, Ludwig, Tsipras, Dimitris, and Vladu, Adrian · 2018
Earlier work this paper cites.
Poison frogs! targeted clean-label poisoning attacks on neural networks
Shafahi, Ali, Huang, W Ronny, Najibi, Mahyar, Suciu, Octavian, Studer, Christoph, Dumitras, Tudor, and Goldstein, Tom · 2018
Earlier work this paper cites.
A new backdoor attack in cnns by training set corruption without label poisoning
Barni, Mauro, Kallas, Kassem, and Tondi, Benedetta · 2019
Earlier work this paper cites.
Universal adversarial attacks on text classifiers
Behjati, Melika, Moosavi-Dezfooli, Seyed-Mohsen, Baghshah, Mahdieh Soleymani, and Frossard, Pascal · 2019
Earlier work this paper cites.
A backdoor attack against lstm-based text classification systems
Dai, Jiazhu, Chen, Chuanshuai, and Li, Yufeng · 2019
Earlier work this paper cites.
On physical adversarial patches for object detection
Lee, Mark and Kolter, Zico · 2019
Earlier work this paper cites.
Fooling automated surveillance cameras: adversarial patches to attack person detection
Thys, Simen, Van Ranst, Wiebe, and Goedemé, Toon · 2019
Earlier work this paper cites.
Label-consistent backdoor attacks
Turner, Alexander, Tsipras, Dimitris, and Madry, Aleksander · 2019
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing nlp
Wallace, Eric, Feng, Shi, Kandpal, Nikhil, Gardner, Matt, and Singh, Sameer · 2019
Earlier work this paper cites.
Neural cleanse: Identifying and mitigating backdoor attacks in neural networks
Wang, Bolun, Yao, Yuanshun, Shan, Shawn, Li, Huiying, Viswanath, Bimal, Zheng, Haitao, and Zhao, Ben Y · 2019
Earlier work this paper cites.
Latent backdoor attacks on deep neural networks
Yao, Yuanshun, Li, Huiying, Zheng, Haitao, and Zhao, Ben Y · 2019
Earlier work this paper cites.
Adversarial framing for image and video classification
Zajac, Michał, Zołna, Konrad, Rostamzadeh, Negar, and Pinheiro, Pedro O · 2019
Earlier work this paper cites.
Transferable clean-label poisoning attacks on deep neural nets
Zhu, Chen, Huang, W Ronny, Li, Hengduo, Taylor, Gavin, Studer, Christoph, and Goldstein, Tom · 2019
Earlier work this paper cites.
Universal adversarial perturbations: A survey
Chaubey, Ashutosh, Agrawal, Nikhil, Barnwal, Kavya, Guliani, Keerat K, and Mehta, Pramod · 2020
Earlier work this paper cites.
Universal adversarial attack on attention and the resulting dataset damagenet
Chen, Sizhe, He, Zhengbao, Sun, Chengjin, Yang, Jie, and Huang, Xiaolin · 2020
Cited alongside, same era.
Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks
Croce, Francesco and Hein, Matthias · 2020
Cited alongside, same era.
Adversarial camouflage: Hiding physical-world attacks with natural styles
Duan, Ranjie, Ma, Xingjun, Wang, Yisen, Bailey, James, Qin, A Kai, and Yang, Yun · 2020
Cited alongside, same era.
Backdooring convolutional neural networks via targeted weight perturbations
Dumford, Jacob and Scheirer, Walter · 2020
Cited alongside, same era.
Backdoor attacks and countermeasures on deep learning: A comprehensive review
Gao, Yansong, Doan, Bao Gia, Zhang, Zhi, Ma, Siqi, Zhang, Jiliang, Fu, Anmin, Nepal, Surya, and Kim, Hyoungshick · 2020
Cited alongside, same era.
Badencoder: Backdoor attacks to pre-trained encoders in self-supervised learning
Jia, Jinyuan, Liu, Yupei, and Gong, Neil Zhenqiang · 2022
Later among the works it cites.
Frequency domain model augmentation for adversarial attack
Long, Yuyang, Zhang, Qilong, Zeng, Boheng, Gao, Lianli, Liu, Xianglong, Zhang, Jian, and Song, Jingkuan · 2022
Later among the works it cites.
Hidden trigger backdoor attack on nlp models via linguistic style manipulation
Pan, Xudong, Zhang, Mi, Sheng, Beina, Zhu, Jiaming, and Yang, Min · 2022
Later among the works it cites.
Towards practical deployment-stage backdoor attack on deep neural networks
Qi, Xiangyu, Xie, Tinghao, Pan, Ruizhe, Zhu, Jifeng, Yang, Yong, and Bu, Kai · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Ramesh, Aditya, Dhariwal, Prafulla, Nichol, Alex, Chu, Casey, and Chen, Mark · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Can adversarial weight perturbations inject neural backdoors
Garg, Siddhant, Kumar, Adarsh, Goel, Vibhor, and Liang, Yingyu · 2020
Cited alongside, same era.
Universal litmus patterns: Revealing backdoor attacks in cnns
Kolouri, Soheil, Saha, Aniruddha, Pirsiavash, Hamed, and Hoffmann, Heiko · 2020
Cited alongside, same era.
Invisible backdoor attacks on deep neural networks via steganography and regularization
Li, Shaofeng, Xue, Minhui, Zhao, Benjamin Zi Hao, Zhu, Haojin, and Zhang, Xinpeng · 2020
Cited alongside, same era.
Composite backdoor attack for deep neural network by mixing existing benign features
Lin, Junyu, Xu, Lei, Liu, Yingqi, and Zhang, Xiangyu · 2020
Cited alongside, same era.
Deep k-nn defense against clean-label data poisoning attacks
Peri, Neehar, Gupta, Neal, Huang, W Ronny, Fowl, Liam, Zhu, Chen, Feizi, Soheil, Goldstein, Tom, and Dickerson, John P · 2020
Cited alongside, same era.
Tbt: Targeted neural network attack with bit trojan
Rakin, Adnan Siraj, He, Zhezhi, and Fan, Deliang · 2020
Cited alongside, same era.
Hidden trigger backdoor attacks
Saha, Aniruddha, Subramanya, Akshayvarun, and Pirsiavash, Hamed · 2020
Cited alongside, same era.
Backdoor attacks on self-supervised learning
Saha, Aniruddha, Tejankar, Ajinkya, Koohpayegani, Soroush Abbasi, and Pirsiavash, Hamed · 2022
Later among the works it cites.
Dynamic backdoor attacks against machine learning models
Salem, Ahmed, Wen, Rui, Backes, Michael, Ma, Shiqing, and Zhang, Yang · 2022
Later among the works it cites.
Dual-key multimodal backdoors for visual question answering
Walmer, Matthew, Sikka, Karan, Sur, Indranil, Shrivastava, Abhinav, and Jha, Susmit · 2022
Later among the works it cites.
An invisible black-box backdoor attack through frequency domain
Wang, Tong, Yao, Yuan, Xu, Feng, An, Shengwei, Tong, Hanghang, and Wang, Ting · 2022
Later among the works it cites.
Badclip: Trigger-aware prompt learning for backdoor attacks on clip
Bai, Jiawang, Gao, Kuofeng, Min, Shaobo, Xia, Shu-Tao, Li, Zhifeng, and Liu, Wei · 2023
Later among the works it cites.
Image hijacks: Adversarial images can control generative models at runtime
Bailey, Luke, Ong, Euan, Russell, Stuart, and Emmons, Scott · 2023
Later among the works it cites.
Cleanclip: Mitigating data poisoning attacks in multimodal contrastive learning
Bansal, Hritik, Singhi, Nishad, Yang, Yu, Yin, Fan, Grover, Aditya, and Chang, Kai-Wei · 2023
Later among the works it cites.
Are aligned neural networks adversarially aligned?
Carlini, Nicholas, Nasr, Milad, Choquette-Choo, Christopher A, Jagielski, Matthew, Gao, Irena, Awadalla, Anas, Koh, Pang Wei, Ippolito, Daphne, Lee, Katherine, Tramer, Florian, et al · 2023
Later among the works it cites.
On the robustness of large multimodal models against image adversarial attacks
Cui, Xuanimng, Aparcedo, Alejandro, Jang, Young Kyun, and Lim, Ser-Nam · 2023
Later among the works it cites.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Dai, Wenliang, Li, Junnan, Li, Dongxu, Tiong, Anthony Meng Huat, Zhao, Junqi, Wang, Weisheng, Li, Boyang, Fung, Pascale, and Hoi, Steven · 2023
Later among the works it cites.
Palm-e: An embodied multimodal language model
Driess, Danny, Xia, Fei, Sajjadi, Mehdi SM, Lynch, Corey, Chowdhery, Aakanksha, Ichter, Brian, Wahid, Ayzaan, Tompson, Jonathan, Vuong, Quan, Yu, Tianhe, et al · 2023
Later among the works it cites.
Scaling laws for adversarial attacks on language model activations
Fort, Stanislav · 2023
Later among the works it cites.
Backdooring multimodal learning
Han, Xingshuo, Wu, Yutong, Zhang, Qingjie, Zhou, Yuan, Xu, Yuan, Qiu, Han, Xu, Guowen, and Zhang, Tianwei · 2023
Later among the works it cites.
Composite backdoor attacks against large language models
Huang, Hai, Zhao, Zhengyu, Backes, Michael, Shen, Yun, and Zhang, Yang · 2023
Later among the works it cites.
Backdoor attacks for in-context learning with language models
Kandpal, Nikhil, Jagielski, Matthew, Tramèr, Florian, and Carlini, Nicholas · 2023
Later among the works it cites.
Badclip: Dual-embedding guided backdoor attack on multimodal contrastive learning
Liang, Siyuan, Zhu, Mingli, Liu, Aishan, Wu, Baoyuan, Cao, Xiaochun, and Chang, Ee-Chien · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Visual adversarial examples jailbreak aligned large language models
Qi, Xiangyu, Huang, Kaixuan, Panda, Ashwinee, Wang, Mengdi, and Mittal, Prateek · 2023
Later among the works it cites.
On the adversarial robustness of multi-modal foundation models
Schlarmann, Christian and Hein, Matthias · 2023
Later among the works it cites.
Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models
Shayegani, Erfan, Dong, Yue, and Abu-Ghazaleh, Nael · 2023
Later among the works it cites.
Tijo: Trigger inversion with joint optimization for defending multimodal backdoored models
Sur, Indranil, Sikka, Karan, Walmer, Matthew, Koneripalli, Kaushik, Roy, Anirban, Lin, Xiao, Divakaran, Ajay, and Jha, Susmit · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, Hugo, Lavril, Thibaut, Izacard, Gautier, Martinet, Xavier, Lachaux, Marie-Anne, Lacroix, Timothée, Rozière, Baptiste, Goyal, Naman, Hambro, Eric, Azhar, Faisal, et al · 2023
Later among the works it cites.
How many unicorns are in this image? a safety evaluation benchmark for vision llms
Tu, Haoqin, Cui, Chenhang, Wang, Zijun, Zhou, Yiyang, Zhao, Bingchen, Han, Junlin, Zhou, Wangchunshu, Yao, Huaxiu, and Xie, Cihang · 2023
Later among the works it cites.
Effective backdoor mitigation depends on the pre-training objective
Verma, Sahil, Bhatt, Gantavya, Schwarzschild, Avi, Singhal, Soumye, Das, Arnav Mohanty, Shah, Chirag, Dickerson, John P, and Bilmes, Jeff · 2023
Later among the works it cites.
Rab: Provable robustness against backdoor attacks
Weber, Maurice, Xu, Xiaojun, Karlaš, Bojan, Zhang, Ce, and Li, Bo · 2023
Later among the works it cites.
Badchain: Backdoor chain-of-thought prompting for large language models
Xiang, Zhen, Jiang, Fengqing, Xiong, Zidi, Ramasubramanian, Bhaskar, Poovendran, Radha, and Li, Bo · 2023
Later among the works it cites.
Narcissus: A practical clean-label backdoor attack with limited information
Zeng, Yi, Pan, Minzhou, Just, Hoang Anh, Lyu, Lingjuan, Qiu, Meikang, and Jia, Ruoxi · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, Deyao, Chen, Jun, Shen, Xiaoqian, Li, Xiang, and Elhoseiny, Mohamed · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Zou, Andy, Wang, Zifan, Kolter, J Zico, and Fredrikson, Matt · 2023
Later among the works it cites.