Fetching the paper…
Reading the bibliography…
Under a commonly-studied backdoor poisoning attack against classification models, an attacker adds a small trigger to a subset of the training data, such that the presence of this trigger at test time causes the classifier to always predict some target class.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Deng Jia, Su Hao, Krause Jonathan, Satheesh Sanjeev, Ma Sean, Huang Zhiheng, Karpathy Andrej, Khosla Aditya, Michael Bernstein, Berg Alexander C., and Fei-Fei Li · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Adversarial machine learning at scale
Alexey Kurakin, Ian Goodfellow, and Samy Bengio · 2016
Earlier work this paper cites.
Why should i trust you?: Explaining the predictions of any classifier
Tulio Ribeiro Marco, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
Tom B. Brown, Dandelion Mane, Aurko Roy, Martín Abadi, and Justin Gilmer · 2017
Earlier work this paper cites.
Targeted backdoor attacks on deep learning systems using data poisoning
Xinyun Chen, Chang Liu, Bo Li, Kimberly Lu, and Dawn Song · 2017
Earlier work this paper cites.
Badnets: Identifying vulnerabilities in the machine learning model supply chain
Tianyu Gu, Dolan-Gavitt Brendan, and Siddharth Garg · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
R Selvarajk Ramprasaath, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra · 2017
Earlier work this paper cites.
Rise: Randomized input sampling for explanation of black-box models
Vitali Petsiuk, Abir Das, and Saenko Saenko · 2018
Earlier work this paper cites.
Spectral signatures in backdoor attacks
Brandon Tran, Jerry Li, and Aleksander Madry · 2018
Earlier work this paper cites.
Certified adversarial robustness via randomized smoothing
Jeremy M Cohen, Elan Rosenfeld, and J. Zico Kolter · 2019
Cited alongside, same era.
Adversarial robustness as a prior for learned representations
Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Brandon Tran, and Aleksander Madry · 2019
Cited alongside, same era.
Strip: A defence against trojan attacks on deep neural networks
Yansong Gao, Chang Xu, Derui Wang, Shiping Chen, Damith C. Ranasinghe, and Surya Nepal · 2019
Cited alongside, same era.
Adversarial examples are not bugs, they are features
Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, and Aleksander Madry · 2019
Cited alongside, same era.
Are perceptually-aligned gradients a general property of robust classifiers?
Simran Kaur, Jeremy Cohen, and Zachary C. Lipton · 2019
Rethinking the trigger of backdoor attack
Yiming Li, Tongqing Zhai, Baoyuan Wu, Yong Jiang, Zhifeng Li, and Shutao Xia · 2020
Closest in time.
Reflection backdoor: A natural backdoor attack on deep neural networks
Yunfei Liu, Xingjun Ma, James Bailey, and Feng Lu · 2020
Closest in time.
Challenge round 0 (dry run) test dataset, 2020
Michael Paul Majurski · 2020
Closest in time.
Input-aware dynamic backdoor attack
Tuan Anh Nguyen and Anh Tran · 2020
Closest in time.
Hidden trigger backdoor attacks
Aniruddha Saha, Akshayvarun Subraymanya, and Pirsiavash Hamed · 2020
Closest in time.
Denoised smoothing: A provable defense for pretrained classifiers
Hadi Salman, Mingjie Sun, Greg Yang, Ashish Kapoor, and J. Zico Kolter · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Invisible backdoor attacks on deep neural networks via steganography and regularization
Shaofeng Li, Minhui Xue, Benjamin Zi Hao Zhao, Haojin Zhu, and Xinpeng Zhang · 2019
Cited alongside, same era.
Defending neural backdoors via generative distribution modeling
Ximing Qiao, Yukun Yang, and Hai Li · 2019
Cited alongside, same era.
Provably robust deep learning via adversarially trained smoothed classifiers
Hadi Salman, Greg Yang, Jerry Li, Pengchuan Zhang, Huan Zhang, Ilya Razenshteyn, and Sebastien Bubeck · 2019
Cited alongside, same era.
Image synthesis with a single (robust) classifier
Shibani Santurkar, Dimitris Tsipras, Brandon Tran, Andrew Ilyas, Logan Engstrom, and Aleksander Madry · 2019
Cited alongside, same era.
Robustness may be at odds with accuracy
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Alexander Turner, and Aleksander Madry · 2019
Cited alongside, same era.
Clean-label backdoor attacks, 2019
Alexander Turner, Dimitris Tsipras, and Aleksander Madry · 2019
Cited alongside, same era.
Neural cleanse: Identifying and mitigating backdoor attacks in neural networks
Bolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li, Bimal Viswanath, Haitao Zheng, and Ben Y. Zhao · 2019
Cited alongside, same era.
Facehack: Triggering backdoored facial recognition systems using facial characteristics
Esha Sarkar, Hadjer Benkraouda, and Michail Maniatakos · 2020
Closest in time.
Exposing backdoors in robust machine learning models
Ezekiel Soremekun, Sakshi Udeshi, and Sudipta Chattopadhyay · 2020
Closest in time.
Practical detection of trojan neural networks: Data-limited and data-free cases
Ren Wang, Gaoyuan Zhang, Sijia Liu, Pin-Yu Chen, Jinjun Xiong, and Meng Wang · 2020
Closest in time.
On the trade-off between adversarial and backdoor robustness
Cheng-Hsin Weng, Yan-Ting Lee, and Shan-Hung (Brandon) Wu · 2020
Closest in time.
Beyond a Gaussian denoiser: Residual learning of deep CNN for image denoising
Kai Zhang, Wangmeng Zuo, Yunjin Chen, Deyu Meng, and Lei Zhang · 2020
Closest in time.
Refit: A unified watermark removal framework for deep learning systems with limited data
Xinyun Chen, Wenxiao Wang, Chris Bender, Yiming Ding, Ruoxi Jia, Bo Li, and Dawn Song · 2021
Closest in time.
Wanet – imperceptible warping-based backdoor attack
Anh Nguyen and Anh Tran · 2021
Closest in time.