Fetching the paper…
Reading the bibliography…
Stability guarantees have emerged as a principled way to evaluate feature attributions, but existing certification methods rely on heavily smoothed classifiers and often produce conservative guarantees.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2013
Earlier work this paper cites.
Analysis of boolean functions
Ryan O’Donnell · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2016
Earlier work this paper cites.
" why should i trust you?" explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
Evaluating the visualization of what a deep neural network has learned
Wojciech Samek, Alexander Binder, Grégoire Montavon, Sebastian Lapuschkin, and Klaus-Robert Müller · 2016
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee · 2017
Earlier work this paper cites.
Smoothgrad: removing noise by adding noise
Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda Viégas, and Martin Wattenberg · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim · 2018
Earlier work this paper cites.
Anchors: High-precision model-agnostic explanations
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2018
Earlier work this paper cites.
Lipschitz regularity of deep neural networks: analysis and efficient estimation
Aladin Virmaux and Kevin Scaman · 2018
Earlier work this paper cites.
Sorting out lipschitz function approximation
Cem Anil, James Lucas, and Roger Grosse · 2019
Earlier work this paper cites.
Machine learning interpretability: A survey on methods and metrics
Diogo V Carvalho, Eduardo M Pereira, and Jaime S Cardoso · 2019
Earlier work this paper cites.
Certified adversarial robustness via randomized smoothing
Jeremy Cohen, Elan Rosenfeld, and Zico Kolter · 2019
Earlier work this paper cites.
Efficient and accurate estimation of lipschitz constants for deep neural networks
Mahyar Fazlyab, Alexander Robey, Hamed Hassani, Manfred Morari, and George Pappas · 2019
Earlier work this paper cites.
Limitations of the lipschitz constant as a defense against adversarial examples
Todd Huster, Cho-Yu Jason Chiang, and Ritu Chadha · 2019
Earlier work this paper cites.
The (un) reliability of saliency methods
Pieter-Jan Kindermans, Sara Hooker, Julius Adebayo, Maximilian Alber, Kristof T Schütt, Sven Dähne, Dumitru Erhan, and Been Kim · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Explanation in artificial intelligence: Insights from the social sciences
Tim Miller · 2019
Earlier work this paper cites.
Towards a rigorous evaluation of xai methods on time series
Udo Schlegel, Hiba Arnout, Mennatallah El-Assady, Daniela Oelke, and Daniel A Keim · 2019
Earlier work this paper cites.
Interpretable and fine-grained visual explanations for convolutional neural networks
Jorg Wagner, Jan Mathias Kohler, Tobias Gindele, Leon Hetzel, Jakob Thaddaus Wiedemer, and Sven Behnke · 2019
Earlier work this paper cites.
On the (in) fidelity and sensitivity of explanations
Chih-Kuan Yeh, Cheng-Yu Hsieh, Arun Suggala, David I Inouye, and Pradeep K Ravikumar · 2019
Earlier work this paper cites.
Influence functions in deep learning are fragile
Samyadeep Basu, Philip Pope, and Soheil Feizi · 2020
Earlier work this paper cites.
Challenging common interpretability assumptions in feature attribution explanations
Jonathan Dinu, Jeffrey Bigham, and J Zico Kolter · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy · 2020
Cited alongside, same era.
Probabilistic lipschitz analysis of neural networks
Ravi Mangal, Kartik Sarangmath, Aditya V Nori, and Alessandro Orso · 2020
Cited alongside, same era.
Fooling lime and shap: Adversarial attacks on post hoc explanation methods
Dylan Slack, Sophie Hilgard, Emily Jia, Sameer Singh, and Himabindu Lakkaraju · 2020
Cited alongside, same era.
The many shapley values for model explanation
Mukund Sundararajan and Amir Najmi · 2020
Cited alongside, same era.
Safe planning in dynamic environments using conformal prediction
Lars Lindemann, Matthew Cleaveland, Gihyun Shim, and George J Pappas · 2023
Later among the works it cites.
Faithful chain-of-thought reasoning
Qing Lyu, Shreya Havaldar, Adam Stein, Li Zhang, Delip Rao, Eric Wong, Marianna Apidianaki, and Chris Callison-Burch · 2023
Later among the works it cites.
From anecdotal evidence to quantitative evaluation methods: A systematic review on evaluating explainable ai
Meike Nauta, Jan Trienes, Shreyasi Pathak, Elisa Nguyen, Michelle Peters, Yasmin Schmitt, Jörg Schlötterer, Maurice Van Keulen, and Christin Seifert · 2023
Later among the works it cites.
Explainable deep learning methods in medical image classification: A survey
Cristiano Patrício, João C Neves, and Luís F Teixeira · 2023
Later among the works it cites.
Hierarchical randomized smoothing
Yan Scholten, Jan Schuchardt, Aleksandar Bojchevski, and Stephan Günnemann · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards global explanations of convolutional neural networks with concept attribution
Weibin Wu, Yuxin Su, Xixian Chen, Shenglin Zhao, Irwin King, Michael R Lyu, and Yu-Wing Tai · 2020
Cited alongside, same era.
Provably efficient, succinct, and precise explanations
Guy Blanc, Jane Lange, and Li-Yang Tan · 2021
Cited alongside, same era.
The out-of-distribution problem in explainability and search methods for feature importance explanations
Peter Hase, Harry Xie, and Mohit Bansal · 2021
Cited alongside, same era.
Improved, deterministic smoothing for l_1 certified robustness
Alexander J Levine and Soheil Feizi · 2021
Cited alongside, same era.
The computational complexity of understanding binary classifier decisions
Stephan Wäldchen, Jan Macdonald, Sascha Hauch, and Gitta Kutyniok · 2021
Cited alongside, same era.
Probabilistic sufficient explanations
Eric Wang, Pasha Khosravi, and Guy Van den Broeck · 2021
Cited alongside, same era.
Post hoc explanations may be ineffective for detecting unknown spurious correlation
Julius Adebayo, Michael Muelly, Harold Abelson, and Been Kim · 2022
Cited alongside, same era.
Anton Xue, Rajeev Alur, and Eric Wong · 2023
Later among the works it cites.
Cpsign: conformal prediction for cheminformatics modeling
Staffan Arvidsson McShane, Ulf Norinder, Jonathan Alvarsson, Ernst Ahlberg, Lars Carlsson, and Ola Spjuth · 2024
Later among the works it cites.
A diagnostic study of explainability techniques for text classification
Pepa Atanasova · 2024
Later among the works it cites.
Impossibility theorems for feature attribution
Blair Bilodeau, Natasha Jaques, Pang Wei Koh, and Been Kim · 2024
Later among the works it cites.
On the evaluation consistency of attribution-based explanations
Jiarui Duan, Haoling Li, Haofei Zhang, Hao Jiang, Mengqi Xue, Li Sun, Mingli Song, and Jie Song · 2024
Later among the works it cites.
Empirical validation of conformal prediction for trustworthy skin lesions classification
Jamil Fayyad, Shadi Alijani, and Homayoun Najjaran · 2024
Later among the works it cites.
Provably stable feature rankings with shap and lime
Jeremy Goldwasser and Giles Hooker · 2024
Later among the works it cites.
The fix benchmark: Extracting features interpretable to experts
Helen Jin, Shreya Havaldar, Chaehyeon Kim, Anton Xue, Weiqiu You, Helen Qu, Marco Gatti, Daniel Hashimoto, Bhuvnesh Jain, Amin Madani, Masao Sako, Lyle Ungar, and Eric Wong · 2024
Later among the works it cites.
Rethinking robustness of model attributions
Sandesh Kamath, Sankalp Mittal, Amit Deshpande, and Vineeth N Balasubramanian · 2024
Later among the works it cites.
Analyzing explainer robustness via probabilistic lipschitzness of prediction functions
Zulqarnain Q Khan, Davin Hill, Aria Masoomi, Joshua T Bone, and Jennifer Dy · 2024
Later among the works it cites.
Toward explainable artificial intelligence for precision pathology
Frederick Klauschen, Jonas Dippel, Philipp Keyl, Philipp Jurmeister, Michael Bockmayr, Andreas Mock, Oliver Buchstab, Maximilian Alber, Lukas Ruff, Grégoire Montavon, et al · 2024
Later among the works it cites.
On the robustness of removal-based feature attributions
Chris Lin, Ian Covert, and Su-In Lee · 2024
Later among the works it cites.
Towards faithful model explanation in nlp: A survey
Qing Lyu, Marianna Apidianaki, and Chris Callison-Burch · 2024
Later among the works it cites.
Explainable reinforcement learning: A survey and comparative review
Stephanie Milani, Nicholay Topin, Manuela Veloso, and Fei Fang · 2024
Later among the works it cites.
A comprehensive and reliable feature attribution method: Double-sided remove and reconstruct (dorar)
Dong Qin, George T Amariucai, Daji Qiao, Yong Guan, and Shen Fu · 2024
Later among the works it cites.
Explainable ai and law: an evidential survey
Karen McGregor Richmond, Satya M Muddamsetty, Thomas Gammeltoft-Hansen, Henrik Palmer Olsen, and Thomas B Moeslund · 2024
Later among the works it cites.
A comprehensive taxonomy for explainable artificial intelligence: a systematic survey of surveys on methods and concepts
Gesina Schwalbe and Bettina Finzel · 2024
Later among the works it cites.
Benchmarking llms via uncertainty quantification
Fanghua Ye, Mingming Yang, Jianhui Pang, Longyue Wang, Derek F Wong, Emine Yilmaz, Shuming Shi, and Zhaopeng Tu · 2024
Later among the works it cites.
Mfaba: A more faithful and accelerated boundary-based attribution method for deep neural networks
Zhiyu Zhu, Huaming Chen, Jiayu Zhang, Xinyi Wang, Zhibo Jin, Minhui Xue, Dongxiao Zhu, and Kim-Kwang Raymond Choo · 2024
Later among the works it cites.
On the computational tractability of the (many) shapley values
Reda Marzouk, Shahaf Bassan, Guy Katz, and Colin de la Higuera · 2025
Closest in time.