Fetching the paper…
Reading the bibliography…
Different users of machine learning methods require different explanations, depending on their goals.
Some remarks on the measurability of certain sets
Paul Erdös · 1945
Earlier work this paper cites.
Introduction to topology
Bert Mendelson · 1990
Earlier work this paper cites.
Topological Spaces: including a treatment of multi-valued functions, vector spaces, and convexity
Claude Berge · 1997
Earlier work this paper cites.
Set-valued analysis
Jean-Pierre Aubin and Hélène Frankowska · 2009
Earlier work this paper cites.
scikit-image: image processing in Python
Stéfan van der Walt, Johannes L. Schönberger, Juan Nunez-Iglesias, François Boulogne, Joshua D. Warner, Neil Yager, Emmanuelle Gouillart, Tony Yu, and the scikit-image contributors · 2011
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2014
Earlier work this paper cites.
Generalizing the poincaré–miranda theorem: the avoiding cones condition
Alessandro Fonda and Paolo Gidoni · 2016
Earlier work this paper cites.
"Why should I trust you?" Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee · 2017
Earlier work this paper cites.
Smoothgrad: removing noise by adding noise
Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda Viégas, and Martin Wattenberg · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
Counterfactual explanations without opening the black box: Automated decisions and the GDPR
Sandra Wachter, Brent Mittelstadt, and Chris Russell · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim · 2018
Earlier work this paper cites.
On the robustness of interpretability methods
David Alvarez-Melis and Tommi S Jaakkola · 2018
Earlier work this paper cites.
Explanations based on the missing: Towards contrastive explanations with pertinent negatives
Amit Dhurandhar, Pin-Yu Chen, Ronny Luss, Chun-Chen Tu, Paishun Ting, Karthikeyan Shanmugam, and Payel Das · 2018
Earlier work this paper cites.
A survey of methods for explaining black box models
Riccardo Guidotti, Anna Monreale, Salvatore Ruggieri, Franco Turini, Fosca Giannotti, and Dino Pedreschi · 2018
Earlier work this paper cites.
The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery
Zachary C Lipton · 2018
Earlier work this paper cites.
Explanations can be manipulated and geometry is to blame
Ann-Kathrin Dombrowski, Maximillian Alber, Christopher Anders, Marcel Ackermann, Klaus-Robert Müller, and Pan Kessel · 2019
Cited alongside, same era.
Interpretation of neural networks is fragile
Amirata Ghorbani, Abubakar Abid, and James Zou · 2019
Cited alongside, same era.
A benchmark for interpretability methods in deep neural networks
Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim · 2019
Cited alongside, same era.
Shalmali Joshi, Oluwasanmi Koyejo, Warut Vijitbenjaronk, Been Kim, and Joydeep Ghosh · 2019
Cited alongside, same era.
The (un) reliability of saliency methods
Pieter-Jan Kindermans, Sara Hooker, Julius Adebayo, Maximilian Alber, Kristof T Schütt, Sven Dähne, Dumitru Erhan, and Been Kim · 2019
Cited alongside, same era.
Fooling LIME and SHAP: Adversarial attacks on post hoc explanation methods
Dylan Slack, Sophie Hilgard, Emily Jia, Sameer Singh, and Himabindu Lakkaraju · 2020
Later among the works it cites.
Counterfactual explanations for machine learning: A review
Sahil Verma, John Dickerson, and Keegan Hines · 2020
Later among the works it cites.
Towards the unification and robustness of perturbation and gradient based explanations
Sushant Agarwal, Shahin Jabbari, Chirag Agarwal, Sohini Upadhyay, Steven Wu, and Himabindu Lakkaraju · 2021
Later among the works it cites.
Geometrically enriched latent spaces
Georgios Arvanitidis, Soren Hauberg, and Bernhard Schölkopf · 2021
Later among the works it cites.
Counterfactual evaluation for explainable ai
Yingqiang Ge, Shuchang Liu, Zelong Li, Shuyuan Xu, Shijie Geng, Yunqi Li, Juntao Tan, Fei Sun, and Yongfeng Zhang · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The dangers of post-hoc interpretability: unjustified counterfactual explanations
Thibault Laugel, Marie-Jeanne Lesot, Christophe Marsala, Xavier Renard, and Marcin Detyniecki · 2019
Cited alongside, same era.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Cynthia Rudin · 2019
Cited alongside, same era.
Explainable AI: Interpreting, Explaining and Visualizing Deep Learning
Wojciech Samek, Grégoire Montavon, Andrea Vedaldi, Lars Kai Hansen, and Klaus-Robert Müller · 2019
Cited alongside, same era.
Actionable recourse in linear classification
Berk Ustun, Alexander Spangher, and Yang Liu · 2019
Cited alongside, same era.
Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI
Alejandro Barredo Arrieta, Natalia Díaz-Rodríguez, Javier Del Ser, Adrien Bennetot, Siham Tabik, Alberto Barbado, Salvador García, Sergio Gil-López, Daniel Molina, Richard Benjamins, et al · 2020
Cited alongside, same era.
Multi-objective counterfactual explanations
Susanne Dandl, Christoph Molnar, Martin Binder, and Bernd Bischl · 2020
Cited alongside, same era.
Opportunities and challenges in Explainable Artificial Intelligence (XAI): A survey
Arun Das and Paul Rad · 2020
Cited alongside, same era.
A survey of algorithmic recourse: contrastive explanations and consequential recommendations
Amir-Hossein Karimi, Gilles Barthe, Bernhard Schölkopf, and Isabel Valera · 2021
Later among the works it cites.
If only we had better counterfactual explanations
Mark T Keane, Eoin M Kenny, Eoin Delaney, and Barry Smyth · 2021
Later among the works it cites.
Towards robust and reliable algorithmic recourse
Sohini Upadhyay, Shalmali Joshi, and Himabindu Lakkaraju · 2021
Later among the works it cites.
Evaluating the quality of machine learning explanations: A survey on methods and metrics
Jianlong Zhou, Amir H Gandomi, Fang Chen, and Andreas Holzinger · 2021
Later among the works it cites.
Impossibility theorems for feature attribution
Blair L. Bilodeau, Natasha Jaques, Pang Wei Koh, and Been Kim · 2022
Closest in time.
Consistent counterfactuals for deep models
Emily Black, Zifan Wang, Matt Fredrikson, and Anupam Datta · 2022
Closest in time.
Post-hoc explanations fail to achieve their purpose in adversarial contexts
Sebastian Bordt, Michèle Finck, Eric Raidl, and Ulrike von Luxburg · 2022
Closest in time.
The robustness of counterfactual explanations over time
Andrea Ferrario and Michele Loi · 2022
Closest in time.
Interpretable Machine Learning
Christoph Molnar · 2022
Closest in time.
On the trade-off between actionable explanations and the right to be forgotten
Martin Pawelczyk, Tobias Leemann, Asia Biega, and Gjergji Kasneci · 2022
Closest in time.
Trustworthy Machine Learning
Kush R. Varshney · 2022
Closest in time.
Robust counterfactual explanations for neural networks with probabilistic guarantees
Faisal Hamman, Erfaun Noorani, Saumitra Mishra, Daniele Magazzeni, and Sanghamitra Dutta · 2023
Closest in time.
Diagnosing ai explanation methods with folk concepts of behavior
Alon Jacovi, Jasmijn Bastings, Sebastian Gehrmann, Yoav Goldberg, and Katja Filippova · 2023
Closest in time.