Fetching the paper…
Reading the bibliography…
Deep neural networks obtain state-of-the-art performance on a series of tasks.
How to explain individual classification decisions
David Baehrens, Timon Schroeter, Stefan Harmeling, Motoaki Kawanabe, Katja Hansen, and Klaus-Robert Müller · 2010
Earlier work this paper cites.
An efficient explanation of individual classifications using game theory
Erik Štrumbelj and Igor Kononenko · 2010
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D Zeiler and Rob Fergus · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy · 2015
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Sebastian Bach, Alexander Binder, Grégoire Montavon, Frederick Klauschen, Klaus-Robert Müller, and Wojciech Samek · 2015
Earlier work this paper cites.
Deepfool: a simple and accurate method to fool deep neural networks
Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, and Pascal Frossard · 2016
Earlier work this paper cites.
The limitations of deep learning in adversarial settings
Nicolas Papernot, Patrick McDaniel, Somesh Jha, Matt Fredrikson, Z Berkay Celik, and Ananthram Swami · 2016
Earlier work this paper cites.
A boundary tilting persepective on the phenomenon of adversarial examples
Thomas Tanay and Lewis Griffin · 2016
Earlier work this paper cites.
Robustness of classifiers: from adversarial to random noise
Alhussein Fawzi, Seyed-Mohsen Moosavi-Dezfooli, and Pascal Frossard · 2016
Earlier work this paper cites.
Why should I trust you?: Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
The mythos of model interpretability
Zachary C Lipton · 2016
Earlier work this paper cites.
Algorithmic transparency via quantitative input influence: Theory and experiments with learning systems
Anupam Datta, Shayak Sen, and Yair Zick · 2016
Earlier work this paper cites.
Distributional smoothing with virtual adversarial training
Takeru Miyato, Shin-ichi Maeda, Masanori Koyama, Ken Nakae, and Shin Ishii · 2016
Earlier work this paper cites.
Distillation as a defense to adversarial perturbations against deep neural networks
Nicolas Papernot, Patrick McDaniel, Xi Wu, Somesh Jha, and Ananthram Swami · 2016
Earlier work this paper cites.
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Attacking machine learning with adversarial examples
Ian Goodfellow, Nicolas Papernot, Sandy Huang, Yan Duan, and Peter Abbeel · 2017
Earlier work this paper cites.
Adversarial machine learning at scale
Alexey Kurakin, Ian Goodfellow, and Samy Bengio · 2017
Earlier work this paper cites.
Towards evaluating the robustness of neural networks
Nicholas Carlini and David Wagner · 2017
Earlier work this paper cites.
Zoo: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models
Pin-Yu Chen, Huan Zhang, Yash Sharma, Jinfeng Yi, and Cho-Jui Hsieh · 2017
Earlier work this paper cites.
Delving into transferable adversarial examples and black-box attacks
Yanpei Liu, Xinyun Chen, Chang Liu, and Dawn Song · 2017
Cited alongside, same era.
Practical black-box attacks against machine learning
Nicolas Papernot, Patrick McDaniel, Ian Goodfellow, Somesh Jha, Z Berkay Celik, and Ananthram Swami · 2017
Cited alongside, same era.
Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje · 2017
Cited alongside, same era.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee · 2017
Cited alongside, same era.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Cited alongside, same era.
Adversarial learning and explainability in structured datasets
Prasad Chalasani, Somesh Jha, Aravind Sadagopan, and Xi Wu · 2018
Later among the works it cites.
Ensemble adversarial training: Attacks and defenses
Florian Tramèr, Alexey Kurakin, Nicolas Papernot, Ian Goodfellow, Dan Boneh, and Patrick McDaniel · 2018
Later among the works it cites.
Pixeldefend: Leveraging generative models to understand and defend against adversarial examples
Yang Song, Taesup Kim, Sebastian Nowozin, Stefano Ermon, and Nate Kushman · 2018
Later among the works it cites.
Feature squeezing: Detecting adversarial examples in deep neural networks
Weilin Xu, David Evans, and Yanjun Qi · 2018
Later among the works it cites.
Towards robust neural networks via random self-ensemble
Xuanqing Liu, Minhao Cheng, Huan Zhang, and Cho-Jui Hsieh · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Amirata Ghorbani, Abubakar Abid, and James Zou · 2017
Cited alongside, same era.
Adversarial examples detection in deep networks with convolutional filter statistics
Xin Li and Fuxin Li · 2017
Cited alongside, same era.
Early methods for detecting adversarial images
Dan Hendrycks and Kevin Gimpel · 2017
Cited alongside, same era.
On the (statistical) detection of adversarial examples
Kathrin Grosse, Praveen Manoharan, Nicolas Papernot, Michael Backes, and Patrick McDaniel · 2017
Cited alongside, same era.
Adversarial and clean data are not twins
Zhitao Gong, Wenlu Wang, and Wei-Shinn Ku · 2017
Cited alongside, same era.
On detecting adversarial perturbations
Jan Hendrik Metzen, Tim Genewein, Volker Fischer, and Bastian Bischoff · 2017
Cited alongside, same era.
Detecting adversarial samples from artifacts
Reuben Feinman, Ryan R Curtin, Saurabh Shintre, and Andrew B Gardner · 2017
Cited alongside, same era.
Provable defenses against adversarial examples via the convex outer adversarial polytope
Eric Wong and J Zico Kolter · 2018
Later among the works it cites.
Training verified learners with learned verifiers
Krishnamurthy Dvijotham, Sven Gowal, Robert Stanforth, Relja Arandjelovic, Brendan O’Donoghue, Jonathan Uesato, and Pushmeet Kohli · 2018
Later among the works it cites.
Adversarially robust generalization requires more data
Ludwig Schmidt, Shibani Santurkar, Dimitris Tsipras, Kunal Talwar, and Aleksander Madry · 2018
Later among the works it cites.
There is no free lunch in adversarial robustness (but there are unexpected benefits)
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Alexander Turner, and Aleksander Madry · 2018
Later among the works it cites.
Enhancing robustness of machine learning systems via data transformations
Arjun Nitin Bhagoji, Daniel Cullina, Chawin Sitawarin, and Prateek Mittal · 2018
Later among the works it cites.
Characterizing adversarial subspaces using local intrinsic dimensionality
Xingjun Ma, Bo Li, Yisen Wang, Sarah M. Erfani, Sudanthi Wijewickrema, Grant Schoenebeck, Michael E. Houle, Dawn Song, and James Bailey · 2018
Later among the works it cites.
A simple unified framework for detecting out-of-distribution samples and adversarial attacks
Kimin Lee, Kibok Lee, Honglak Lee, and Jinwoo Shin · 2018
Later among the works it cites.
Attacks meet interpretability: Attribute-steered detection of adversarial samples
Guanhong Tao, Shiqing Ma, Yingqi Liu, and Xiangyu Zhang · 2018
Later among the works it cites.
Detecting adversarial perturbations with saliency
Chiliang Zhang, Zuochang Ye, Yan Wang, and Zhimou Yang · 2018
Later among the works it cites.
Prior convictions: Black-box adversarial attacks with bandits and priors
Andrew Ilyas, Logan Engstrom, and Aleksander Madry · 2019
Closest in time.
L-shapley and C-shapley: Efficient model interpretation for structured data
Jianbo Chen, Le Song, Martin J. Wainwright, and Michael I. Jordan · 2019
Closest in time.
How sensitive are sensitivity-based explanations?
Chih-Kuan Yeh, Cheng-Yu Hsieh, Arun Sai Suggala, David Inouye, and Pradeep Ravikumar · 2019
Closest in time.
Rob-gan: Generator, discriminator, and adversarial attacker
Xuanqing Liu and Cho-Jui Hsieh · 2019
Closest in time.
Certified robustness to adversarial examples with differential privacy
Mathias Lecuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu, and Suman Jana · 2019
Closest in time.
Adv-bnn: Improved adversarial defense through robust bayesian neural network
Xuanqing Liu, Yao Li, Chongruo Wu, and Cho-Jui Hsieh · 2019
Closest in time.