Fetching the paper…
Reading the bibliography…
Vision Transformers (ViTs) have become prominent models for solving various vision tasks.
A. Gretton, K. Borgwardt, M. Rasch, B. Schölkopf, and A. Smola, “A kernel method for the two-sample-problem,” NeurIPS , vol. 19, 2006
2006
Earlier work this paper cites.
A. Gretton, K. M. Borgwardt, M. J. Rasch, B. Schölkopf, and A. Smola, “A kernel two-sample test,” The Journal of Machine Learning Research , vol. 13, no. 1, pp. 723–773, 2012
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
S. Bach, A. Binder, G. Montavon, F. Klauschen, K.-R. Müller, and W. Samek, “On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation,” PloS one , vol. 10, no. 7, p. e0130140, 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
M. T. Ribeiro and C. Guestrin, “” why should i trust you?” explaining the predictions of any classifier,” in Proceedings of the 22nd ACM SIGKDD , 2016, pp. 1135–1144
2016
Earlier work this paper cites.
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba, “Learning deep features for discriminative localization,” in Proceedings of CVPR , 2016, pp. 2921–2929
2016
Earlier work this paper cites.
A. Vaswani et al. , “Attention is all you need,” NeurIPS , vol. 30, 2017
2017
Earlier work this paper cites.
J. Kim and J. Canny, “Interpretable learning for self-driving cars by visualizing causal attention,” in ICCV , 2017, pp. 2942–2950
2017
Earlier work this paper cites.
S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” NeurIPS , vol. 30, 2017
2017
Earlier work this paper cites.
A. Shrikumar, P. Greenside, and A. Kundaje, “Learning important features through propagating activation differences,” in ICML . PMLR, 2017, pp. 3145–3153
2017
Earlier work this paper cites.
M. Sundararajan, A. Taly, and Q. Yan, “Axiomatic attribution for deep networks,” in ICML . PMLR, 2017, pp. 3319–3328
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in ICCV , 2017, pp. 618–626
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. Adebayo, J. Gilmer, M. Muelly, I. Goodfellow, M. Hardt, and B. Kim, “Sanity checks for saliency maps,” NeurIPS , vol. 31, 2018
2018
Earlier work this paper cites.
M. Wu, M. Hughes, S. Parbhoo, M. Zazzi, V. Roth, and F. Doshi-Velez, “Beyond sparsity: Tree regularization of deep models for interpretability,” in AAAI , vol. 32, no. 1, 2018
2018
Earlier work this paper cites.
Q. Zhang, Y. N. Wu, and S.-C. Zhu, “Interpretable convolutional neural networks,” in Proceedings of CVPR , 2018, pp. 8827–8836
2018
Earlier work this paper cites.
W. Samek, G. Montavon, A. Vedaldi, L. K. Hansen, and K.-R. Müller, Explainable AI: interpreting, explaining and visualizing deep learning . Springer Nature, 2019, vol. 11700
2019
Earlier work this paper cites.
P.-J. Kindermans, S. Hooker, J. Adebayo, M. Alber, K. T. Schütt, S. Dähne, D. Erhan, and B. Kim, “The (un) reliability of saliency methods,” in Explainable AI: Interpreting, Explaining and Visualizing Deep Learning . Springer, 2019, pp. 267–280
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
R. Fong, M. Patrick, and A. Vedaldi, “Understanding deep networks via extremal perturbations and smooth masks,” in ICCV , 2019
2019
Earlier work this paper cites.
C. Chen et al. , “This looks like that: deep learning for interpretable image recognition,” NeurIPS , vol. 32, 2019
2019
Cited alongside, same era.
S. Serrano and N. A. Smith, “Is attention interpretable?” arXiv preprint arXiv:1906.03731 , 2019
2019
Cited alongside, same era.
2020
Cited alongside, same era.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in ECCV . Springer, 2020, pp. 213–229
2020
Cited alongside, same era.
2021
Later among the works it cites.
H. Touvron, M. Cord, A. Sablayrolles, G. Synnaeve, and H. Jégou, “Going deeper with image transformers,” in ICCV , 2021, pp. 32–42
2021
Later among the works it cites.
W. Gao, F. Wan, X. Pan, Z. Peng, Q. Tian, Z. Han, B. Zhou, and Q. Ye, “Ts-cam: Token semantic coupled attention map for weakly supervised object localization,” in ICCV , 2021, pp. 2886–2895
2021
Later among the works it cites.
T. Yuan, X. Li, H. Xiong, H. Cao, and D. Dou, “Explaining information flow inside vision transformers using markov chain,” in eXplainable AI approaches for debugging and diagnosis. , 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Stiglic, P. Kocbek, N. Fijacko, M. Zitnik, K. Verbert, and L. Cilar, “Interpretability of machine learning-based prediction models in healthcare,” Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery , vol. 10, no. 5, p. e1379, 2020
2020
Cited alongside, same era.
A. B. Arrieta et al. , “Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai,” Information fusion , vol. 58, pp. 82–115, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
M. Al-Shedivat, A. Dubey, and E. P. Xing, “Contextual explanation networks.” J. Mach. Learn. Res. , vol. 21, pp. 194–1, 2020
2020
Cited alongside, same era.
Z. Chen, Y. Bei, and C. Rudin, “Concept whitening for interpretable image recognition,” Nature Machine Intelligence , 2020
2020
Cited alongside, same era.
P. Angelov and E. Soares, “Towards explainable deep neural networks (xdnn),” Neural Networks , vol. 130, pp. 185–194, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
P. Sturmfels, S. Lundberg, and S.-I. Lee, “Visualizing the impact of feature attribution baselines,” Distill , vol. 5, no. 1, p. e22, 2020
2020
Cited alongside, same era.
B. Pan, R. Panda, Y. Jiang, Z. Wang, R. Feris, and A. Oliva, “Ia-red: Interpretability-aware redundancy reduction for vision transformers,” NeurIPS , vol. 34, pp. 24 898–24 911, 2021
2021
Later among the works it cites.
R. Alharbi, M. N. Vu, and M. T. Thai, “Learning interpretation with explainable knowledge distillation,” in Big Data . IEEE, 2021, pp. 705–714
2021
Later among the works it cites.
O. Barkan, E. Hauon, A. Caciularu, O. Katz, I. Malkiel, O. Armstrong, and N. Koenigstein, “Grad-sam: Explaining transformers via gradient self-attention maps,” in Proceedings of CIKM , 2021, pp. 2882–2887
2021
Later among the works it cites.
H. Shah, P. Jain, and P. Netrapalli, “Do input gradients highlight discriminative features?” NeurIPS , vol. 34, pp. 2046–2059, 2021
2021
Later among the works it cites.
H. Xu, Z. Cai, and W. Li, “Privacy-preserving mechanisms for multi-label image recognition,” ACM TKDD , vol. 16, no. 4, pp. 1–21, 2022
2022
Later among the works it cites.
R. Wang, D. Chen, Z. Wu, Y. Chen, X. Dai, M. Liu, Y.-G. Jiang, L. Zhou, and L. Yuan, “Bevt: Bert pretraining of video transformers,” in CVPR , 2022, pp. 14 733–14 743
2022
Later among the works it cites.
Z. Liu, J. Ning, Y. Cao, Y. Wei, Z. Zhang, S. Lin, and H. Hu, “Video swin transformer,” in CVPR , 2022, pp. 3202–3211
2022
Later among the works it cites.
Y. Qiang, C. Li, M. Brocanelli, and D. Zhu, “Counterfactual interpolation augmentation (cia): A unified approach to enhance fairness and explainability of dnn,” in IJCAI , 2022, pp. 732–739
2022
Later among the works it cites.
Y. Qiang, D. Pan, C. Li, X. Li, R. Jang, and D. Zhu, “Attcat: Explaining transformers via attentive class activation tokens,” in NeurIPS , 2022
2022
Later among the works it cites.
S. Kim, J. Nam, and B. C. Ko, “Vit-net: Interpretable vision transformers with neural tree decoder,” in ICML . PMLR, 2022, pp. 11 162–11 172
2022
Later among the works it cites.
2022
Later among the works it cites.
M. Böhle, M. Fritz, and B. Schiele, “B-cos networks: alignment is all we need for interpretability,” in CVPR , 2022, pp. 10 329–10 338
2022
Later among the works it cites.
J. Guo, K. Han, H. Wu, Y. Tang, X. Chen, Y. Wang, and C. Xu, “Cmt: Convolutional neural networks meet vision transformers,” in CVPR , 2022, pp. 12 175–12 185
2022
Later among the works it cites.
Z. Chen, C. Wang, Y. Wang, G. Jiang, Y. Shen, Y. Tai, C. Wang, W. Zhang, and L. Cao, “Lctr: On awakening the local continuity of transformer for weakly supervised object localization,” in AAAI , vol. 36, no. 1, 2022, pp. 410–418
2022
Later among the works it cites.
S. Gupta, S. Lakhotia, A. Rawat, and R. Tallamraju, “Vitol: Vision transformer for weakly supervised object localization,” in CVPR , 2022, pp. 4101–4110
2022
Later among the works it cites.
C. Li, H. Bagher-Ebadian, V. Goddla, I. J. Chetty, and D. Zhu, “Focalunetr: A focal transformer for boundary-aware segmentation of ct images,” MICCAI , 2023
2023
Closest in time.
X. Li, D. Pan, C. Li, Y. Qiang, and D. Zhu, “Negative flux aggregation to estimate feature attributions,” in IJCAI , 2023, pp. 446–454
2023
Closest in time.
X. Li, X. Li, D. Pan, Y. Qiang, and D. Zhu, “Learning compact features via in-training representation alignment,” in Proceedings of AAAI , vol. 37, no. 7, 2023, pp. 8675–8683
2023
Closest in time.
Y. Qiang, C. Li, P. Khanduri, and D. Zhu, “Fairness-aware vision transformer via debiased self-attention,” in European Conference on Computer Vision . Springer, 2024, pp. 358–376
2024
Closest in time.
2024
Closest in time.
C. Li, P. Khanduri, Y. Qiang, R. I. Sultan, I. Chetty, and D. Zhu, “Auto-prompting sam for mobile friendly 3d medical image segmentation,” WACV , 2025
2025
Closest in time.