Fetching the paper…
Reading the bibliography…
Current methods for the interpretability of discriminative deep neural networks commonly rely on the model's input-gradients, i.e., the gradients of the output logits w.r.t.
Probabilistic interpretation of feedforward classification network outputs, with relationships to statistical pattern recognition
John S Bridle · 1990
Earlier work this paper cites.
A stochastic estimator of the trace of the influence matrix for laplacian smoothing splines
Michael F Hutchinson · 1990
Earlier work this paper cites.
Fast exact multiplication by the hessian
Barak A Pearlmutter · 1994
Earlier work this paper cites.
Estimation of non-normalized statistical models by score matching
Aapo Hyvärinen · 2005
Earlier work this paper cites.
Regularized estimation of image statistics by score matching
Durk P Kingma and Yann LeCun · 2010
Earlier work this paper cites.
Randomized algorithms for estimating the trace of an implicit symmetric positive semi-definite matrix
Haim Avron and Sivan Toledo · 2011
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
Pascal Vincent · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Max Welling and Yee W Teh · 2011
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2013
Earlier work this paper cites.
Inceptionism: Going deeper into neural networks
Alexander Mordvintsev, Christopher Olah, and Mike Tyka · 2015
Earlier work this paper cites.
Visualizing deep convolutional neural networks using natural pre-images
Aravindh Mahendran and Andrea Vedaldi · 2016
Earlier work this paper cites.
Synthesizing the preferred inputs for neurons in neural networks via deep generator networks
Anh Nguyen, Alexey Dosovitskiy, Jason Yosinski, Thomas Brox, and Jeff Clune · 2016
Earlier work this paper cites.
Evaluating the visualization of what a deep neural network has learned
Wojciech Samek, Alexander Binder, Grégoire Montavon, Sebastian Lapuschkin, and Klaus-Robert Müller · 2016
Earlier work this paper cites.
Andrew Slavin Ross and Finale Doshi-Velez · 2017
Cited alongside, same era.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra · 2017
Cited alongside, same era.
Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje · 2017
Cited alongside, same era.
Smoothgrad: removing noise by adding noise
Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda Viégas, and Martin Wattenberg · 2017
Cited alongside, same era.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim · 2018
Cited alongside, same era.
Interpretation of neural networks is fragile
Amirata Ghorbani, Abubakar Abid, and James Zou · 2019
Later among the works it cites.
Fooling neural network interpretations via adversarial model manipulation
Juyeon Heo, Sunghwan Joo, and Taesup Moon · 2019
Later among the works it cites.
Are perceptually-aligned gradients a general property of robust classifiers?
Simran Kaur, Jeremy Cohen, and Zachary C Lipton · 2019
Later among the works it cites.
Image synthesis with a single (robust) classifier
Shibani Santurkar, Andrew Ilyas, Dimitris Tsipras, Logan Engstrom, Brandon Tran, and Aleksander Madry · 2019
Later among the works it cites.
Generative modeling by estimating gradients of the data distribution
Yang Song and Stefano Ermon · 2019
Later among the works it cites.
Sliced score matching: A scalable approach to density and score estimation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards better understanding of gradient-based attribution methods for deep neural networks
Marco Ancona, Enea Ceolini, Cengiz Oztireli, and Markus Gross · 2018
Cited alongside, same era.
Concise explanations of neural networks using adversarial training
Prasad Chalasani, Jiefeng Chen, Amrita Roy Chowdhury, Somesh Jha, and Xi Wu · 2018
Cited alongside, same era.
Improving dnn robustness to adversarial attacks using jacobian regularization
Daniel Jakubovitz and Raja Giryes · 2018
Cited alongside, same era.
How good is my gan?
Konstantin Shmelkov, Cordelia Schmid, and Karteek Alahari · 2018
Cited alongside, same era.
Explanations can be manipulated and geometry is to blame
Ann-Kathrin Dombrowski, Maximillian Alber, Christopher Anders, Marcel Ackermann, Klaus-Robert Müller, and Pan Kessel · 2019
Cited alongside, same era.
Adversarial robustness as a prior for learned representations
Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Brandon Tran, and Aleksander Madry · 2019
Cited alongside, same era.
On the connection between adversarial robustness and saliency map interpretability
Christian Etmann, Sebastian Lunz, Peter Maass, and Carola-Bibiane Schönlieb · 2019
Cited alongside, same era.
Yang Song, Sahaj Garg, Jiaxin Shi, and Stefano Ermon · 2019
Later among the works it cites.
Full-gradient representation for neural network visualization
Suraj Srinivas and François Fleuret · 2019
Later among the works it cites.
Fooling network interpretation in image classification
Akshayvarun Subramanya, Vipin Pillai, and Hamed Pirsiavash · 2019
Later among the works it cites.
Implicit gradient regularization
David GT Barrett and Benoit Dherin · 2020
Closest in time.
Your classifier is secretly an energy based model and you should treat it like one
Will Grathwohl, Kuan-Chieh Wang, Joern-Henrik Jacobsen, David Duvenaud, Mohammad Norouzi, and Kevin Swersky · 2020
Closest in time.
Efficient learning of generative models via finite-difference score matching
Tianyu Pang, Taufik Xu, Chongxuan Li, Yang Song, Stefano Ermon, and Jun Zhu · 2020
Closest in time.
Interpretable deep learning under fire
Xinyang Zhang, Ningfei Wang, Hua Shen, Shouling Ji, Xiapu Luo, and Ting Wang · 2020
Closest in time.