Fetching the paper…
Reading the bibliography…
We show new connections between adversarial learning and explainability for deep neural networks (DNNs).
On the (in)fidelity and sensitivity of explanations
Yeh, C.-K., Hsieh, C.-Y., Suggala, A., Inouye, D. I., and Ravikumar, P. K · 1901
Earlier work this paper cites.
Adversarial examples are not bugs, they are features
Ilyas, A., Santurkar, S., Tsipras, D., Engstrom, L., Tran, B., and Madry, A · 1905
Earlier work this paper cites.
One explanation does not fit all: A toolkit and taxonomy of AI explainability techniques
Arya, V., Bellamy, R. K. E., Chen, P.-Y., Dhurandhar, A., Hind, M., Hoffman, S. C., Houde, S., Vera Liao, Q., Luss, R., Mojsilović, A., Mourad, S., Pedemonte, P., Raghavendra, R., Richards, J., Sattigeri, P., Shanmugam, K., Singh, M., Varshney, K. R., Wei, D., and Zhang, Y · 1909
Earlier work this paper cites.
Comparing measures of sparsity
Hurley, N. and Rickard, S · 2009
Earlier work this paper cites.
Robustness and regularization of support vector machines
Xu, H., Caramanis, C., and Mannor, S · 2009
Earlier work this paper cites.
How to explain individual classification decisions
Baehrens, D., Schroeter, T., Harmeling, S., Kawanabe, M., Hansen, K., and Müller, K.-R · 2010
Earlier work this paper cites.
MNIST handwritten digit database
LeCun, Y. and Cortes, C · 2010
Earlier work this paper cites.
Evasion attacks against machine learning at test time
Biggio, B., Corona, I., Maiorca, D., Nelson, B., Šrndić, N., Laskov, P., Giacinto, G., and Roli, F · 2013
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Simonyan, K., Vedaldi, A., and Zisserman, A · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R · 2013
Earlier work this paper cites.
Minimax sparse logistic regression for very high-dimensional feature selection
Tan, M., Tsang, I. W., and Wang, L · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Goodfellow, I. J., Shlens, J., and Szegedy, C · 2014
Cited alongside, same era.
Towards ultrahigh dimensional feature selection for big data
Tan, M., Tsang, I. W., and Wang, L · 2014
Cited alongside, same era.
The limitations of deep learning in adversarial settings
Papernot, N., McDaniel, P., Jha, S., Fredrikson, M., Berkay Celik, Z., and Swami, A · 2015
Cited alongside, same era.
Shaham, U., Yamada, Y., and Negahban, S · 2015
Cited alongside, same era.
“why should I trust you?”: Explaining the predictions of any classifier
Regularizing deep networks using efficient layerwise adversarial training
Sankaranarayanan, S., Jain, A., Chellappa, R., and Lim, S. N · 2017
Later among the works it cites.
Axiomatic attribution for deep networks
Sundararajan, M., Taly, A., and Yan, Q · 2017
Later among the works it cites.
Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms
Xiao, H., Rasul, K., and Vollgraf, R · 2017
Later among the works it cites.
Adversarial examples: Attacks and defenses for deep learning
Yuan, X., He, P., Zhu, Q., and Li, X · 2017
Later among the works it cites.
On the robustness of interpretability methods
Alvarez-Melis, D. and Jaakkola, T. S · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ribeiro, M. T., Singh, S., and Guestrin, C · 2016
Cited alongside, same era.
High-dimensional classification by sparse logistic regression
Abramovich, F. and Grinshtein, V · 2017
Cited alongside, same era.
Towards better understanding of gradient-based attribution methods for deep neural networks
Ancona, M., Ceolini, E., Öztireli, C., and Gross, M · 2017
Cited alongside, same era.
UCI machine learning repository, 2017
Dheeru, D. and Karra Taniskidou, E · 2017
Cited alongside, same era.
What the success of brain imaging implies about the neural code
Guest, O. and Love, B. C · 2017
Cited alongside, same era.
A unified approach to interpreting model predictions
Lundberg, S. M. and Lee, S.-I · 2017
Cited alongside, same era.
Towards deep learning models resistant to adversarial attacks
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A · 2017
Cited alongside, same era.
Closest in time.
Certifying some distributional robustness with principled adversarial training
Sinha, A., Namkoong, H., and Duchi, J · 2018
Closest in time.
A new angle on l2 regularization
Tanay, T. and Griffin, L. D · 2018
Closest in time.
Robustness may be at odds with accuracy
Tsipras, D., Santurkar, S., Engstrom, L., Turner, A., and Madry, A · 2018
Closest in time.
Bridging adversarial robustness and gradient interpretability
Kim, B., Seo, J., and Jeon, T · 2019
Closest in time.
Interpretable Machine Learning
Molnar, C · 2019
Closest in time.
Does interpretability of neural networks imply adversarial robustness?
Noack, A., Ahern, I., Dou, D., and Li, B · 2019
Closest in time.