Fetching the paper…
Reading the bibliography…
How do neural networks extract patterns from pixels? Feature visualizations attempt to answer this important question by visualizing highly activating patterns through optimization.
A simple saliency method that passes the sanity checks
Gupta, A. and Arora, S · 1905
Earlier work this paper cites.
How to manipulate CNNs to make them lie: the GradCAM case
Viering, T., Wang, Z., Loog, M., and Eisemann, E · 1907
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Hornik, K., Stinchcombe, M., and White, H · 1989
Earlier work this paper cites.
Toward trustworthy AI development: mechanisms for supporting verifiable claims
Brundage, M., Avin, S., Wang, J., Belfield, H., Krueger, G., Hadfield, G., Khlaaf, H., Yang, J., Toner, H., Fong, R., et al · 2004
Earlier work this paper cites.
On the optimality of conditional expectation as a Bregman predictor
Banerjee, A., Guo, X., and Wang, H · 2005
Earlier work this paper cites.
Visualizing higher-layer features of a deep network
Erhan, D., Bengio, Y., Courville, A., and Vincent, P · 2009
Earlier work this paper cites.
Towards falsifiable interpretability research
Leavitt, M. L. and Morcos, A · 2010
Earlier work this paper cites.
Thinking in circuits: toward neurobiological explanation in cognitive neuroscience
Pulvermüller, F., Garagnani, M., and Wennekers, T · 2014
Earlier work this paper cites.
Striving for simplicity: The all convolutional net
Springenberg, J. T., Dosovitskiy, A., Brox, T., and Riedmiller, M · 2014
Earlier work this paper cites.
Deep neural networks: a new framework for modeling biological vision and brain information processing
Kriegeskorte, N · 2015
Earlier work this paper cites.
Understanding deep image representations by inverting them
Mahendran, A. and Vedaldi, A · 2015
Earlier work this paper cites.
DeepDream-a code example for visualizing neural networks
Mordvintsev, A., Olah, C., and Tyka, M · 2015
Earlier work this paper cites.
Adversarial manipulation of deep representations
Sabour, S., Cao, Y., Faghri, F., and Fleet, D. J · 2015
Earlier work this paper cites.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S. E., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2015
Earlier work this paper cites.
Understanding neural networks through deep visualization
Yosinski, J., Clune, J., Nguyen, A., Fuchs, T., and Lipson, H · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Visualizing deep convolutional neural networks using natural pre-images
Mahendran, A. and Vedaldi, A · 2016
Earlier work this paper cites.
Nguyen, A., Yosinski, J., and Clune, J · 2016
Earlier work this paper cites.
”why should i trust you?” explaining the predictions of any classifier
Ribeiro, M. T., Singh, S., and Guestrin, C · 2016
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations
Bau, D., Zhou, B., Khosla, A., Oliva, A., and Torralba, A · 2017
Earlier work this paper cites.
Targeted backdoor attacks on deep learning systems using data poisoning
Chen, X., Liu, C., Li, B., Lu, K., and Song, D · 2017
Earlier work this paper cites.
Badnets: Identifying vulnerabilities in the machine learning model supply chain
Gu, T., Dolan-Gavitt, B., and Garg, S · 2017
Earlier work this paper cites.
Formal guarantees on the robustness of a classifier against adversarial manipulation
Hein, M. and Andriushchenko, M · 2017
Cited alongside, same era.
A unified approach to interpreting model predictions
Lundberg, S. M. and Lee, S.-I · 2017
Cited alongside, same era.
Feature visualization
Olah, C., Mordvintsev, A., and Schubert, L · 2017
Cited alongside, same era.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D · 2017
Cited alongside, same era.
Axiomatic attribution for deep networks
Sundararajan, M., Taly, A., and Yan, Q · 2017
Cited alongside, same era.
Sanity checks for saliency maps
Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I., Hardt, M., and Kim, B · 2018
Cited alongside, same era.
Do adversarially robust imagenet models transfer better?
Salman, H., Ilyas, A., Engstrom, L., Kapoor, A., and Madry, A · 2020
Later among the works it cites.
Fooling LIME and SHAP: Adversarial attacks on post hoc explanation methods
Slack, D., Hilgard, S., Jia, E., Singh, S., and Lakkaraju, H · 2020
Later among the works it cites.
On adaptive attacks to adversarial example defenses
Tramer, F., Carlini, N., Brendel, W., and Madry, A · 2020
Later among the works it cites.
Exemplary natural images explain CNN activations better than state-of-the-art feature visualization
Borowski, J., Zimmermann, R. S., Schepers, J., Geirhos, R., Wallis, T. S. A., Bethge, M., and Brendel, W · 2021
Later among the works it cites.
How well do feature visualizations support causal understanding of CNN activations?
Zimmermann, R. S., Borowski, J., Geirhos, R., Bethge, M., Wallis, T. S. A., and Brendel, W · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Turning your weakness into a strength: Watermarking deep neural networks by backdooring
Adi, Y., Baum, C., Cisse, M., Pinkas, B., and Keshet, J · 2018
Cited alongside, same era.
Understanding deep neural networks with rectified linear units
Arora, R., Basu, A., Mianjy, P., and Mukherjee, A · 2018
Cited alongside, same era.
A theoretical explanation for perplexing behaviors of backpropagation-based visualizations
Nie, W., Zhang, Y., and Patel, A · 2018
Cited alongside, same era.
The building blocks of interpretability
Olah, C., Satyanarayan, A., Johnson, I., Carter, S., Schubert, L., Ye, K., and Mordvintsev, A · 2018
Cited alongside, same era.
Fairwashing: the risk of rationalization
Aïvodji, U., Arai, H., Fortineau, O., Gambs, S., Hara, S., and Tapp, A · 2019
Cited alongside, same era.
Neural population control via deep image synthesis
Bashivan, P., Kar, K., and DiCarlo, J. J · 2019
Cited alongside, same era.
Chen, K., Garudadri, H., and Rao, B. D · 2022
Later among the works it cites.
Backdoor attacks on the DNN interpretation system
Fang, S. and Choromanska, A · 2022
Later among the works it cites.
Attribution-based explanations that provide recourse cannot be robust
Fokkema, H., de Heide, R., and van Erven, T · 2022
Later among the works it cites.
Dataset security for machine learning: Data poisoning, backdoor attacks, and defenses
Goldblum, M., Tsipras, D., Xie, C., Chen, X., Schwarzschild, A., Song, D., Madry, A., Li, B., and Goldstein, T · 2022
Later among the works it cites.
Which explanation should I choose? A function approximation perspective to characterizing post hoc explanations
Han, T., Srinivas, S., and Lakkaraju, H · 2022
Later among the works it cites.
Backdooring explainable machine learning
Noppel, M., Peter, L., and Wressnegger, C · 2022
Later among the works it cites.
Towards better understanding attribution methods
Rao, S., Böhle, M., and Schiele, B · 2022
Later among the works it cites.
Washing the unwashable : On the (im)possibility of fairwashing detection
Shamsabadi, A. S., Yaghini, M., Dullerud, N., Wyllie, S. C., Aïvodji, U., Alaagib, A., Gambs, S., and Papernot, N · 2022
Later among the works it cites.
Circumventing interpretability: How to defeat mind-readers
Sharkey, L · 2022
Later among the works it cites.
Fooling partial dependence via data poisoning
Baniecki, H., Kretowicz, W., and Biecek, P · 2023
Closest in time.
In or out? fixing imagenet out-of-distribution detection evaluation
Bitterwolf, J., Müller, M., and Hein, M · 2023
Closest in time.
Unlocking feature visualization for deeper networks with MAgnitude Constrained Optimization
Fel, T., Boissin, T., Boutin, V., Picard, A., Novello, P., Colin, J., Linsley, D., Rousseau, T., Cadène, R., Gardes, L., et al · 2023
Closest in time.
Pause giant AI experiments: An open letter, 2023
Future of Life Institute · 2023
Closest in time.
Manipulating feature visualizations with gradient slingshots
Bareeva, D., Höhne, M. M.-C., Warnecke, A., Pirch, L., Müller, K.-R., Rieck, K., and Bykov, K · 2024
Closest in time.
Impossibility theorems for feature attribution
Bilodeau, B., Jaques, N., Koh, P. W., and Kim, B · 2024
Closest in time.
Adversarial attacks on the interpretation of neuron activation maximization
Nanfack, G., Fulleringer, A., Marty, J., Eickenberg, M., and Belilovsky, E · 2024
Closest in time.
Measuring mechanistic interpretability at scale without humans
Zimmermann, R. S., Klindt, D. A., and Brendel, W · 2024
Closest in time.