Fetching the paper…
Reading the bibliography…
Concept-based explanations translate the internal representations of deep learning models into a language that humans are familiar with: concepts.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
The caltech-ucsd birds-200-2011 dataset
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie · 2011
Earlier work this paper cites.
Microsoft COCO: common objects in context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, Lubomir D. Bourdev, Ross B. Girshick, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Neural Network with Unbounded Activation Functions is Universal Approximator
Sho Sonoda and Noboru Murata · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick · 2016
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio · 2017
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations
David Bau, Bolei Zhou, Aditya Khosla, Aude Oliva, and Antonio Torralba · 2017
Earlier work this paper cites.
Skin lesion analysis toward melanoma detection: A challenge at the 2017 international symposium on biomedical imaging (isbi), hosted by the international skin imaging collaboration (isic)
Noel C. F. Codella, David Gutman, M. Emre Celebi, Brian Helba, Michael A. Marchetti, Stephen W. Dusza, Aadi Kalloo, Konstantinos Liopyris, Nabin Mishra, Harald Kittler, and Allan Halpern · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim · 2017
Earlier work this paper cites.
dsprites: Disentanglement testing sprites dataset
Loic Matthey, Irina Higgins, Demis Hassabis, and Alexander Lerchner · 2017
Earlier work this paper cites.
Inceptionism: Going deeper into neural networks
Alexander Mordvintsev, Christopher Olah, and Mike Tyka · 2017
Earlier work this paper cites.
Feature visualization
Chris Olah, Alexander Mordvintsev, and Ludwig Schubert · 2017
Earlier work this paper cites.
Places: A 10 million Image Database for Scene Recognition
Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim · 2018
Earlier work this paper cites.
3d shapes dataset
Chris Burgess and Hyunjik Kim · 2018
Earlier work this paper cites.
Net2vec: Quantifying and explaining how concepts are encoded by filters in deep neural networks
Ruth Fong and Andrea Vedaldi · 2018
Earlier work this paper cites.
Seven-point checklist and skin lesion classification using multitask multimodal neural nets
Jeremy Kawahara, Sara Daneshvar, Giuseppe Argenziano, and Ghassan Hamarneh · 2018
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al · 2018
Earlier work this paper cites.
The ham10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions
P. Tschandl, C. Rosendahl, and Kittler H · 2018
Cited alongside, same era.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz · 2018
Cited alongside, same era.
Interpretable basis decomposition for visual explanation
Bolei Zhou, Yiyou Sun, David Bau, and Antonio Torralba · 2018
Cited alongside, same era.
Bcn20000: Dermoscopic lesions in the wild
Marc Combalia, Noel C. F. Codella, Veronica Rotemberg, Brian Helba, Veronica Vilaplana, Ofer Reiter, Allan C. Halpern, Susana Puig, and Josep Malvehy · 2019
Cited alongside, same era.
Towards automatic concept-based explanations
Amirata Ghorbani, James Wexler, James Zou, and Been Kim · 2019
Cited alongside, same era.
BIM: towards quantitative evaluation of interpretability methods with ground truth
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
Trivialaugment: Tuning-free yet state-of-the-art data augmentation
Samuel G. Müller and Frank Hutter · 2021
Later among the works it cites.
Best of both worlds: local and global explanations with human-understandable concepts
Jessica Schrouff, Sebastien Baur, Shaobo Hou, Diana Mincu, Eric Loreaux, Ralph Blanes, James Wexler, Alan Karthikesalingam, and Been Kim · 2021
Later among the works it cites.
From "where" to "what": Towards human-understandable explanations through concept relevance propagation, 2022
Reduan Achtibat, Maximilian Dreyer, Ilona Eisenbraun, Sebastian Bosse, Thomas Wiegand, Wojciech Samek, and Sebastian Lapuschkin · 2022
Later among the works it cites.
Concept gradient: Concept-based interpretation without linear assumption, 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mengjiao Yang and Been Kim · 2019
Cited alongside, same era.
Cutmix: Regularization strategy to train strong classifiers with localizable features
S. Yun, D. Han, S. Chun, S. Oh, Y. Yoo, and J. Choe · 2019
Cited alongside, same era.
Concept whitening for interpretable image recognition
Zhi Chen, Yijie Bei, and Cynthia Rudin · 2020
Cited alongside, same era.
Concept attribution: Explaining CNN decisions to physicians
M. Graziani, V. Andrearczyk, Marchand Maillet S., and H. Müller · 2020
Cited alongside, same era.
What shapes feature representations? Exploring datasets, architectures, and training
Katherine Hermann and Andrew Lampinen · 2020
Cited alongside, same era.
Augment Your Batch: Improving Generalization Through Instance Repetition
Elad Hoffer, Tal Ben-Nun, Itay Hubara, Niv Giladi, Torsten Hoefler, and Daniel Soudry · 2020
Cited alongside, same era.
On interpretability of deep learning based skin lesion classifiers using concept activation vectors
Adriano Lucieri, Muhammad Naseer Bajwa, Stephan Alexander Braun, Muhammad Imran Malik, Andreas Dengel, and Sheraz Ahmed · 2020
Cited alongside, same era.
Andrew Bai, Chih-Kuan Yeh, Pradeep Ravikumar, Neil Y. C. Lin, and Cho-Jui Hsieh · 2022
Later among the works it cites.
What i cannot predict, i do not understand: A human-centered evaluation framework for explainability methods
Julien Colin, Thomas FEL, Remi Cadene, and Thomas Serre · 2022
Later among the works it cites.
Concept activation regions: A generalized framework for concept-based explanations
Jonathan Crabbé and Mihaela van der Schaar · 2022
Later among the works it cites.
Toy models of superposition, 2022
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah · 2022
Later among the works it cites.
Identifying Phenotypic Concepts Discriminating Molecular Breast Cancer Sub-Types
Christoph Fürböck, Matthias Perkonigg, Thomas Helbich, Katja Pinker, Valeria Romeo, and Georg Langs · 2022
Later among the works it cites.
Acquisition of chess knowledge in alphazero
Thomas McGrath, Andrei Kapishnikov, Nenad Tomašev, Adam Pearce, Martin Wattenberg, Demis Hassabis, Been Kim, Ulrich Paquet, and Vladimir Kramnik · 2022
Later among the works it cites.
Craft: Concept recursive activation factorization for explainability
Thomas Fel, Agustin Picard, Louis Bethune, Thibaut Boissin, David Vigouroux, Julien Colin, Rémi Cadène, and Thomas Serre · 2023
Later among the works it cites.
Dividing and conquering a blackbox to a mixture of interpretable models: Route, interpret, repeat
Shantanu Ghosh, Ke Yu, Forough Arabshahi, and Kayhan Batmanghelich · 2023
Later among the works it cites.
Concept Correlation and Its Effects on Concept-Based Models
Lena Heidemann, Maureen Monnet, and Karsten Roscher · 2023
Later among the works it cites.
Emergent world representations: Exploring a sequence model trained on a synthetic task, 2023
Kenneth Li, Aspen K. Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2023
Later among the works it cites.
Do concept bottleneck models obey locality?
Naveen Raman, Mateo Espinosa Zarlenga, Juyeon Heo, and Mateja Jamnik · 2023
Later among the works it cites.
Overlooked factors in concept-based explanations: Dataset choice, concept learnability, and human capability
Vikram V. Ramaswamy, Sunnie S. Y. Kim, Ruth C. Fong, and Olga Russakovsky · 2023
Later among the works it cites.
Towards trustable skin cancer diagnosis via rewriting model’s decision
S. Yan, Z. Yu, X. Zhang, D. Mahapatra, S. S. Chandra, M. Janda, P. Soyer, and Z. Ge · 2023
Later among the works it cites.
Post-hoc concept bottleneck models
Mert Yuksekgonul, Maggie Wang, and James Zou · 2023
Later among the works it cites.
Towards Robust Metrics for Concept Representation Evaluation
Mateo Espinosa Zarlenga, Pietro Barbiero, Zohreh Shams, Dmitry Kazhdan, Umang Bhatt, Adrian Weller, and Mateja Jamnik · 2023
Later among the works it cites.
Understanding inter-concept relationships in concept-based models
Naveen Janaki Raman, Mateo Espinosa Zarlenga, and Mateja Jamnik · 2024
Closest in time.