Fetching the paper…
Reading the bibliography…
In recent years, concept-based approaches have emerged as some of the most promising explainability methods to help us interpret the decisions of Artificial Neural Networks (ANNs).
The approximation of one matrix by another of lower rank
Carl Eckart and Gale Young · 1936
Earlier work this paper cites.
On the abstract properties of linear dependence
Hassler Whitney · 1992
Earlier work this paper cites.
Stability and generalization
Olivier Bousquet and André Elisseeff · 2002
Earlier work this paper cites.
On the equivalence of nonnegative matrix factorization and spectral clustering
Chris Ding, Xiaofeng He, and Horst D Simon · 2005
Earlier work this paper cites.
Concept possession, experimental semantics, and hybrid theories of reference
James Genone and Tania Lombrozo · 2012
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2013
Earlier work this paper cites.
Sparse modeling for image and vision processing
Julien Mairal, Francis Bach, Jean Ponce, et al · 2014
Earlier work this paper cites.
K-sparse autoencoders
Alireza Makhzani and Brendan Frey · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim · 2017
Earlier work this paper cites.
Smoothgrad: removing noise by adding noise
Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda Viégas, and Martin Wattenberg · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra · 2017
Earlier work this paper cites.
Interpretable explanations of black boxes by meaningful perturbation
Ruth C. Fong and Andrea Vedaldi · 2017
Earlier work this paper cites.
Interpretation of neural networks is fragile
Amirata Ghorbani, Abubakar Abid, and James Zou · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim · 2018
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al · 2018
Earlier work this paper cites.
Rise: Randomized input sampling for explanation of black-box models
Vitali Petsiuk, Abir Das, and Kate Saenko · 2018
Earlier work this paper cites.
Efficient neural network robustness certification with general activation functions
Huan Zhang, Tsui-Wei Weng, Pin-Yu Chen, Cho-Jui Hsieh, and Luca Daniel · 2018
Cited alongside, same era.
Mobilenetv2: Inverted residuals and linear bottlenecks
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen · 2018
Cited alongside, same era.
Umap: Uniform manifold approximation and projection for dimension reduction
Leland McInnes, John Healy, and James Melville · 2018
Cited alongside, same era.
Towards better understanding of gradient-based attribution methods for deep neural networks
Marco Ancona, Enea Ceolini, Cengiz Öztireli, and Markus Gross · 2018
Cited alongside, same era.
Dictionary learning algorithms and applications
Bogdan Dumitrescu and Paul Irofti · 2018
Cited alongside, same era.
Sharpening local interpretable model-agnostic explanations for histopathology: improved understandability and reliability
Mara Graziani, Iam Palatnik de Sousa, Marley MBR Vellasco, Eduardo Costa da Silva, Henning Müller, and Vincent Andrearczyk · 2021
Later among the works it cites.
Reliable post hoc explanations: Modeling uncertainty in explainability
Dylan Slack, Anna Hilgard, Sameer Singh, and Himabindu Lakkaraju · 2021
Later among the works it cites.
What i cannot predict, i do not understand: A human-centered evaluation framework for explainability methods
Julien Colin, Thomas Fel, Rémi Cadène, and Thomas Serre · 2021
Later among the works it cites.
The effectiveness of feature attribution methods and its correlation with automatic evaluation scores
Giang Nguyen, Daeyoung Kim, and Anh Nguyen · 2021
Later among the works it cites.
Invertible concept-based explanations for cnn models with non-negative concept activation vectors
Ruihan Zhang, Prashan Madumal, Tim Miller, Krista A Ehinger, and Benjamin IP Rubinstein · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards automatic concept-based explanations
Amirata Ghorbani, James Wexler, James Y Zou, and Been Kim · 2019
Cited alongside, same era.
The (un) reliability of saliency methods
Pieter-Jan Kindermans, Sara Hooker, Julius Adebayo, Maximilian Alber, Kristof T Schütt, Sven Dähne, Dumitru Erhan, and Been Kim · 2019
Cited alongside, same era.
A benchmark for interpretability methods in deep neural networks
Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim · 2019
Cited alongside, same era.
Computing linear restrictions of neural networks
Matthew Sotoudeh and Aditya V. Thakur · 2019
Cited alongside, same era.
Is deep learning ready to satisfy industry needs?
Paolo Tripicchio and Salvatore D’Avella · 2020
Cited alongside, same era.
When explanations lie: Why many modified bp attributions fail
Leon Sixt, Maximilian Granz, and Tim Landgraf · 2020
Cited alongside, same era.
Evaluating explainable ai: Which algorithmic explanations help users predict model behavior?
Peter Hase and Mohit Bansal · 2020
Cited alongside, same era.
Later among the works it cites.
On baselines for local feature attributions
Johannes Haug, Stefan Zürn, Peter El-Jiz, and Gjergji Kasneci · 2021
Later among the works it cites.
Evaluations and methods for explanation through robustness analysis
Cheng-Yu Hsieh, Chih-Kuan Yeh, Xuanqing Liu, Pradeep Ravikumar, Seungyeon Kim, Sanjiv Kumar, and Cho-Jui Hsieh · 2021
Later among the works it cites.
Making sense of dependence: Efficient black-box explanations using dependence measure
Paul Novello, Thomas Fel, and David Vigouroux · 2022
Later among the works it cites.
Towards better understanding attribution methods
Sukrut Rao, Moritz Böhle, and Bernt Schiele · 2022
Later among the works it cites.
HIVE: Evaluating the human interpretability of visual explanations
Sunnie S. Y. Kim, Nicole Meister, Vikram V. Ramaswamy, Ruth Fong, and Olga Russakovsky · 2022
Later among the works it cites.
Do users benefit from interpretable vision? a user study, baseline, and dataset
Leon Sixt, Martin Schuessler, Oana-Iuliana Popescu, Philipp Weiß, and Tim Landgraf · 2022
Later among the works it cites.
Listen to interpret: Post-hoc interpretability for audio networks with nmf
Jayneel Parekh, Sanjeel Parekh, Pavlo Mozharovskyi, Florence d’Alché Buc, and Gaël Richard · 2022
Later among the works it cites.
Out-of-distribution detection with deep nearest neighbors
Yiyou Sun, Yifei Ming, Xiaojin Zhu, and Yixuan Li · 2022
Later among the works it cites.
On the coalitional decomposition of parameters of interest, 2023
Marouane Il Idrissi, Nicolas Bousquet, Fabrice Gamboa, Bertrand Iooss, and Jean-Michel Loubes · 2023
Closest in time.
Craft: Concept recursive activation factorization for explainability
Thomas Fel, Agustin Picard, Louis Bethune, Thibaut Boissin, David Vigouroux, Julien Colin, Rémi Cadène, and Thomas Serre · 2023
Closest in time.
Concept discovery and dataset exploration with singular value decomposition
Mara Graziani, An-phi Nguyen, Laura O’Mahony, Henning Müller, and Vincent Andrearczyk · 2023
Closest in time.
Multi-dimensional concept discovery (mcd): A unifying framework with completeness guarantees
Johanna Vielhaben, Stefan Blücher, and Nils Strodthoff · 2023
Closest in time.
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah · 2023
Closest in time.
Initialization for non-negative matrix factorization: a comprehensive review
Sajad Fathi Hafshejani and Zahra Moaberfard · 2023
Closest in time.