Fetching the paper…
Reading the bibliography…
Post-hoc explanation methods are used with the intent of providing insights about neural networks and are sometimes said to help engender trust in their outputs.
Improving generalization performance using double backpropagation
Harris Drucker and Yann Le Cun · 1992
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy · 2015
Earlier work this paper cites.
Big data’s disparate impact
Solon Barocas and Andrew D Selbst · 2016
Earlier work this paper cites.
Explainable artificial intelligence (xai) darpa-baa-16-53
D Gunning · 2016
Earlier work this paper cites.
Learning deep features for discriminative localization
Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba · 2016
Earlier work this paper cites.
Towards evaluating the robustness of neural networks
Nicholas Carlini and David Wagner · 2017
Earlier work this paper cites.
European Union regulations on algorithmic decision-making and a “right to explanation”
Bryce Goodman and Seth Flaxman · 2017
Earlier work this paper cites.
Adversarial example defense: Ensembles of weak defenses are not strong
Warren He, James Wei, Xinyun Chen, Nicholas Carlini, and Dawn Song · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra · 2017
Earlier work this paper cites.
Smoothgrad: removing noise by adding noise
Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda Viégas, and Martin Wattenberg · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
Evaluating robustness of neural networks with mixed integer programming
Vincent Tjeng, Kai Xiao, and Russ Tedrake · 2017
Earlier work this paper cites.
Visualizing deep neural network decisions: Prediction difference analysis
Luisa M Zintgraf, Taco S Cohen, Tameem Adel, and Max Welling · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim · 2018
Earlier work this paper cites.
Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples
Anish Athalye, Nicholas Carlini, and David Wagner · 2018
Earlier work this paper cites.
Ai2: Safety and robustness certification of neural networks with abstract interpretation
Timon Gehr, Matthew Mirman, Dana Drachsler-Cohen, Petar Tsankov, Swarat Chaudhuri, and Martin Vechev · 2018
Earlier work this paper cites.
On the effectiveness of interval bound propagation for training verifiably robust models
Sven Gowal, Krishnamurthy Dvijotham, Robert Stanforth, Rudy Bunel, Chongli Qin, Jonathan Uesato, Relja Arandjelovic, Timothy Mann, and Pushmeet Kohli · 2018
Earlier work this paper cites.
The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery
Zachary C Lipton · 2018
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2018
Cited alongside, same era.
Differentiable abstract interpretation for provably robust neural networks
Matthew Mirman, Timon Gehr, and Martin Vechev · 2018
Cited alongside, same era.
Anchors: High-precision model-agnostic explanations
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2018
Cited alongside, same era.
Towards fast computation of certified robustness for relu networks
Tsui-Wei Weng, Huan Zhang, Hongge Chen, Zhao Song, Cho-Jui Hsieh, Duane Boning, Inderjit S Dhillon, and Luca Daniel · 2018
Cited alongside, same era.
Feature-guided black-box safety testing of deep neural networks
Matthew Wicker, Xiaowei Huang, and Marta Kwiatkowska · 2018
Cited alongside, same era.
“How do I fool you?” manipulating user trust via misleading black box explanations
Himabindu Lakkaraju and Osbert Bastani · 2020
Later among the works it cites.
Robust and stable black box explanations
Himabindu Lakkaraju, Nino Arsov, and Osbert Bastani · 2020
Later among the works it cites.
A simple defense against adversarial attacks on heatmap explanations
Laura Rieger and Lars Kai Hansen · 2020
Later among the works it cites.
Attributional robustness training using input-gradient spatial alignment
Mayank Singh, Nupur Kumari, Puneet Mangla, Abhishek Sinha, Vineeth N Balasubramanian, and Balaji Krishnamurthy · 2020
Later among the works it cites.
Fooling lime and shap: Adversarial attacks on post hoc explanation methods
Dylan Slack, Sophie Hilgard, Emily Jia, Sameer Singh, and Himabindu Lakkaraju · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jiefeng Chen, Xi Wu, Vaibhav Rastogi, Yingyu Liang, and Somesh Jha · 2019
Cited alongside, same era.
Explanations can be manipulated and geometry is to blame
Ann-Kathrin Dombrowski, Maximillian Alber, Christopher Anders, Marcel Ackermann, Klaus-Robert Müller, and Pan Kessel · 2019
Cited alongside, same era.
Mahyar Fazlyab, Manfred Morari, and George J Pappas · 2019
Cited alongside, same era.
Interpretation of neural networks is fragile
Amirata Ghorbani, Abubakar Abid, and James Zou · 2019
Cited alongside, same era.
Fooling neural network interpretations via adversarial model manipulation
Juyeon Heo, Sunghwan Joo, and Taesup Moon · 2019
Cited alongside, same era.
Abduction-based explanations for machine learning models
Alexey Ignatiev, Nina Narodytska, and Joao Marques-Silva · 2019
Cited alongside, same era.
Automatic model monitoring for data streams
Fábio Pinto, Marco OP Sampaio, and Pedro Bizarro · 2019
Cited alongside, same era.
Zifan Wang, Haofan Wang, Shakul Ramkumar, Matt Fredrikson, Piotr Mardziel, and Anupam Datta · 2020
Later among the works it cites.
Probabilistic safety for bayesian neural networks
Matthew Wicker, Luca Laurenti, Andrea Patane, and Marta Kwiatkowska · 2020
Later among the works it cites.
A game-based approximate verification of deep neural networks with provable guarantees
Min Wu, Matthew Wicker, Wenjie Ruan, Xiaowei Huang, and Marta Kwiatkowska · 2020
Later among the works it cites.
Make sure you’re unsure: A framework for verifying probabilistic specifications
Leonard Berrada, Sumanth Dathathri, Krishnamurthy Dvijotham, Robert Stanforth, Rudy R Bunel, Jonathan Uesato, Sven Gowal, and M Pawan Kumar · 2021
Later among the works it cites.
Provably efficient, succinct, and precise explanations
Guy Blanc, Jane Lange, and Li-Yang Tan · 2021
Later among the works it cites.
On guaranteed optimal robust explanations for nlp models
Emanuele La Malfa, Rhiannon Michelmore, Agnieszka M. Zbrzezny, Nicola Paoletti, and Marta Kwiatkowska · 2021
Later among the works it cites.
Scaling guarantees for nearest counterfactual explanations
Kiarash Mohammadi, Amir-Hossein Karimi, Gilles Barthe, and Isabel Valera · 2021
Later among the works it cites.
A simple and effective method to defend against saliency map attack
Nianwen Si, Heyu Chang, and Yichen Li · 2021
Later among the works it cites.
Adversarial robustness of Bayesian neural networks
Matthew Wicker · 2021
Later among the works it cites.
Bayesian inference with certifiable adversarial robustness
Matthew Wicker, Luca Laurenti, Andrea Patane, Zhuotong Chen, Zheng Zhang, and Marta Kwiatkowska · 2021
Later among the works it cites.
Individual fairness guarantees for neural networks
Elias Benussi, Andrea Patane’, Matthew Wicker, Luca Laurenti, and Marta Kwiatkowska · 2022
Closest in time.
A query-optimal algorithm for finding counterfactuals
Guy Blanc, Caleb Koch, Jane Lange, and Li-Yang Tan · 2022
Closest in time.
Towards robust explanations for deep neural networks
Ann-Kathrin Dombrowski, Christopher J Anders, Klaus-Robert Müller, and Pan Kessel · 2022
Closest in time.