Fetching the paper…
Reading the bibliography…
Concept Bottleneck Models (CBMs) have recently been proposed to address the 'black-box' problem of deep neural networks, by first mapping images to a human-understandable concept space and then linearly combining concepts for classification.
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: ImageNet: A Large-Scale Hierarchical Image Database. In: CVPR. pp. 248–255 (2009)
2009
Earlier work this paper cites.
Krizhevsky, A., Hinton, G., et al.: Learning Multiple Layers of Features from Tiny Images. Technical Report, Computer Science Department, University of Toronto (2009)
2009
Earlier work this paper cites.
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: ImageNet: A Large-Scale Hierarchical Image Database. In: CVPR. pp. 248–255 (2009)
2009
Earlier work this paper cites.
Krizhevsky, A., Hinton, G., et al.: Learning Multiple Layers of Features from Tiny Images. Technical Report, Computer Science Department, University of Toronto (2009)
2009
Earlier work this paper cites.
Wah, C., Branson, S., Welinder, P., Perona, P., Belongie, S.: The Caltech-UCSD Birds-200-2011 Dataset. Tech. Rep. CNS-TR-2011-001, California Institute of Technology (2011)
2011
Earlier work this paper cites.
Patterson, G., Xu, C., Su, H., Hays, J.: The SUN Attribute Database: Beyond Categories for Deeper Scene Understanding. IJCV 108
2014
Earlier work this paper cites.
Patterson, G., Xu, C., Su, H., Hays, J.: The SUN Attribute Database: Beyond Categories for Deeper Scene Understanding. IJCV 108
2014
Earlier work this paper cites.
Bach, S., Binder, A., Montavon, G., Klauschen, F., Müller, K.R., Samek, W.: On Pixel-wise Explanations for Non-Linear Classifier Decisions by Layer-wise Relevance Propagation. PloS one 10
2015
Earlier work this paper cites.
Kingma, D.P., Ba, J.: Adam: A Method for Stochastic Optimization. In: ICLR (2015)
2015
Earlier work this paper cites.
He, K., Zhang, X., Ren, S., Sun, J.: Deep Residual Learning for Image Recognition. In: CVPR. pp. 770–778 (2016)
2016
Earlier work this paper cites.
Hendricks, L.A., Akata, Z., Rohrbach, M., Donahue, J., Schiele, B., Darrell, T.: Generating Visual Explanations. In: ECCV. pp. 3–19. Springer (2016)
2016
Earlier work this paper cites.
Ribeiro, M.T., Singh, S., Guestrin, C.: "Why Should I Trust You?" Explaining the Predictions of any Classifier. In: KDD. pp. 1135–1144 (2016)
2016
Earlier work this paper cites.
He, K., Zhang, X., Ren, S., Sun, J.: Deep Residual Learning for Image Recognition. In: CVPR. pp. 770–778 (2016)
2016
Earlier work this paper cites.
Bau, D., Zhou, B., Khosla, A., Oliva, A., Torralba, A.: Network Dissection: Quantifying Interpretability of Deep Visual Representations. In: CVPR. pp. 6541–6549 (2017)
2017
Earlier work this paper cites.
Lundberg, S.M., Lee, S.I.: A Unified Approach to Interpreting Model Predictions. NeurIPS 30
2017
Earlier work this paper cites.
Ross, A.S., Hughes, M.C., Doshi-Velez, F.: Right for the Right Reasons: Training Differentiable Models by Constraining their Explanations. In: IJCAI. pp. 2662–2670 (2017)
2017
Earlier work this paper cites.
Selvaraju, R.R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., Batra, D.: Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. In: ICCV. pp. 618–626 (2017)
2017
Earlier work this paper cites.
Shrikumar, A., Greenside, P., Kundaje, A.: Learning Important Features through Propagating Activation Differences. In: ICML. pp. 3145–3153 (2017)
2017
Earlier work this paper cites.
Sundararajan, M., Taly, A., Yan, Q.: Axiomatic Attribution for Deep Metworks. In: ICML. pp. 3319–3328. PMLR (2017)
2017
Earlier work this paper cites.
Zhou, B., Lapedriza, A., Khosla, A., Oliva, A., Torralba, A.: Places: A 10 million Image Database for Scene Recognition. IEEE TPAMI (2017)
2017
Earlier work this paper cites.
Zhou, B., Lapedriza, A., Khosla, A., Oliva, A., Torralba, A.: Places: A 10 million Image Database for Scene Recognition. IEEE TPAMI (2017)
2017
Earlier work this paper cites.
Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I., Hardt, M., Kim, B.: Sanity Checks for Saliency Maps. In: NeurIPS. vol. 31 (2018)
2018
Earlier work this paper cites.
Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., Viegas, F., et al.: Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV). In: ICML. pp. 2668–2677 (2018)
2018
Earlier work this paper cites.
Sharma, P., Ding, N., Goodman, S., Soricut, R.: Conceptual Captions: A Cleaned, Hypernymed, Image Alt-Text Dataset for Automatic Image Captioning. In: ACL. pp. 2556–2565 (2018)
2018
Earlier work this paper cites.
Zhou, B., Sun, Y., Bau, D., Torralba, A.: Interpretable Basis Decomposition for Visual Explanation. In: ECCV. pp. 119–134 (2018)
2018
Earlier work this paper cites.
Sharma, P., Ding, N., Goodman, S., Soricut, R.: Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset for Automatic Image Captioning. In: ACL. pp. 2556–2565 (2018)
2018
Earlier work this paper cites.
Chen, C., Li, O., Tao, D., Barnett, A., Rudin, C., Su, J.K.: This Looks Like That: Deep Learning for Interpretable Image Recognition. In: NeurIPS. vol. 32 (2019)
2019
Earlier work this paper cites.
Ghorbani, A., Wexler, J., Zou, J.Y., Kim, B.: Towards Automatic Concept-Based Explanations. In: NeurIPS. vol. 32 (2019)
2019
Earlier work this paper cites.
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., Chintala, S.: PyTorch: An Imperative Style, High-Performance Deep Learning Library. In: NeurIPS (2019)
2019
Earlier work this paper cites.
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.: Language Models are Few-Shot Learners. In: NeurIPS. vol. 33, pp. 1877–1901 (2020)
2020
Cited alongside, same era.
Koh, P.W., Nguyen, T., Tang, Y.S., Mussmann, S., Pierson, E., Kim, B., Liang, P.: Concept Bottleneck Models. In: ICML. pp. 5338–5348 (2020)
2020
Cited alongside, same era.
Sagawa, S., Koh, P.W., Hashimoto, T.B., Liang, P.: Distributionally Robust Neural Networks. In: ICLR (2020)
2020
Cited alongside, same era.
Kokhlikyan, N., Miglani, V., Martin, M., Wang, E., Alsallakh, B., Reynolds, J., Melnikov, A., Kliushkina, N., Araya, C., Yan, S., Reblitz-Richardson, O.: Captum: A Unified and Generic Model Interpretability Library for PyTorch (2020)
2020
Cited alongside, same era.
Sagawa, S., Koh, P.W., Hashimoto, T.B., Liang, P.: Distributionally Robust Neural Networks. In: ICLR (2020)
Fel, T., Picard, A., Bethune, L., Boissin, T., Vigouroux, D., Colin, J., Cadène, R., Serre, T.: CRAFT: Concept Recursive Activation FacTorization for Explainability. In: CVPR. pp. 2711–2721 (2023)
2023
Later among the works it cites.
Graziani, M., Nguyen, A.p., O’Mahony, L., Müller, H., Andrearczyk, V.: Concept Discovery and Dataset Exploration with Singular Value Decomposition. In: ICLRW (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Kim, E., Jung, D., Park, S., Kim, S., Yoon, S.: Probabilistic Concept Bottleneck Models. In: ICML (2023)
2023
Later among the works it cites.
Menon, S., Vondrick, C.: Visual Classification via Description from Large Language Models. In: ICLR (2023)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
Adebayo, J., Muelly, M., Abelson, H., Kim, B.: Post Hoc Explanations may be Ineffective for Detecting Unknown Spurious Correlation. In: ICLR (2021)
2021
Cited alongside, same era.
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., Houlsby, N.: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In: ICLR (2021)
2021
Cited alongside, same era.
Margeloiu, A., Ashman, M., Bhatt, U., Chen, Y., Jamnik, M., Weller, A.: Do Concept Bottleneck Models Learn as Intended? In: ICLRW (2021)
2021
Cited alongside, same era.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning Transferable Visual Models from Natural Language Supervision. In: ICML. pp. 8748–8763 (2021)
2021
Cited alongside, same era.
Zhang, R., Madumal, P., Miller, T., Ehinger, K.A., Rubinstein, B.I.: Invertible Concept-based Explanations for CNN Models with Non-Negative Concept Activation Vectors. In: AAAI. pp. 11682–11690 (2021)
2021
Cited alongside, same era.
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., Houlsby, N.: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In: ICLR (2021)
2021
Cited alongside, same era.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning Transferable Visual Models from Natural Language Supervision. In: ICML. pp. 8748–8763 (2021)
2021
Cited alongside, same era.
2023
Later among the works it cites.
Moayeri, M., Rezaei, K., Sanjabi, M., Feizi, S.: Text-to-Concept (and Back) via Cross-Model Alignment. In: ICML. pp. 25037–25060 (2023)
2023
Later among the works it cites.
Oikarinen, T., Das, S., Nguyen, L.M., Weng, T.W.: Label-Free Concept Bottleneck Models. In: ICLR (2023)
2023
Later among the works it cites.
Oikarinen, T., Weng, T.W.: CLIP-Dissect: Automatic Description of Neuron Representations in Deep Vision Networks. In: ICLR (2023)
2023
Later among the works it cites.
O’Mahony, L., Andrearczyk, V., Müller, H., Graziani, M.: Disentangling Neuron Representations with Concept Vectors. In: CVPRW. pp. 3769–3774 (2023)
2023
Later among the works it cites.
Panousis, K.P., Chatzis, S.: DISCOVER: Making Vision Networks Interpretable via Competition and Dissection. In: NeurIPS (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Panousis, K.P., Ienco, D., Marcos, D.: Sparse Linear Concept Discovery Models. In: ICCVW. pp. 2767–2771 (2023)
2023
Later among the works it cites.
Rao, S., Böhle, M., Parchami-Araghi, A., Schiele, B.: Studying How to Efficiently and Effectively Guide Models with Explanations. In: ICCV. pp. 1922–1933 (2023)
2023
Later among the works it cites.
Roth, K., Kim, J.M., Koepke, A.S., Vinyals, O., Schmid, C., Akata, Z.: Waffling Around for Performance: Visual Classification with Random Words and Broad Concepts. In: ICCV. pp. 15746–15757 (2023)
2023
Later among the works it cites.
Yang, Y., Panagopoulou, A., Zhou, S., Jin, D., Callison-Burch, C., Yatskar, M.: Language in a Bottle: Language Model Guided Concept Bottlenecks for Interpretable Image Classification. In: CVPR. pp. 19187–19197 (2023)
2023
Later among the works it cites.
Yuksekgonul, M., Wang, M., Zou, J.: Post-hoc Concept Bottleneck Models. In: ICLR (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., McLean, B., Burke, J.E., Hume, T., Carter, S., Henighan, T., Olah, C.: Towards Monosemanticity: Decomposing Language Models With Dictionary Learning. Transformer Circuits Thread (2023)
2023
Later among the works it cites.
Cooney, A.: Sparse Autoencoder Library. https://github.com/ai-safety-foundation/sparse_autoencoder (2023)
2023
Later among the works it cites.
Menon, S., Vondrick, C.: Visual Classification via Description from Large Language Models. In: ICLR (2023)
2023
Later among the works it cites.
Oikarinen, T., Das, S., Nguyen, L.M., Weng, T.W.: Label-Free Concept Bottleneck Models. In: ICLR (2023)
2023
Later among the works it cites.
Oikarinen, T., Weng, T.W.: CLIP-Dissect: Automatic Description of Neuron Representations in Deep Vision Networks. In: ICLR (2023)
2023
Later among the works it cites.
Panousis, K.P., Ienco, D., Marcos, D.: Sparse Linear Concept Discovery Models. In: ICCVW. pp. 2767–2771 (2023)
2023
Later among the works it cites.
Ramaswamy, V.V., Kim, S.S., Fong, R., Russakovsky, O.: Overlooked Factors in Concept-Based Explanations: Dataset Choice, Concept Learnability, and Human Capability. In: CVPR. pp. 10932–10941 (2023)
2023
Later among the works it cites.
Rao, S., Böhle, M., Parchami-Araghi, A., Schiele, B.: Studying How to Efficiently and Effectively Guide Models with Explanations. In: ICCV. pp. 1922–1933 (2023)
2023
Later among the works it cites.
Yang, Y., Panagopoulou, A., Zhou, S., Jin, D., Callison-Burch, C., Yatskar, M.: Language in a Bottle: Language Model Guided Concept Bottlenecks for Interpretable Image Classification. In: CVPR. pp. 19187–19197 (2023)
2023
Later among the works it cites.
2024
Closest in time.
Fry, H.: Towards Multimodal Interpretability: Learning Sparse Interpretable Features in Vision Transformers (2024), https://www.lesswrong.com/posts/bCtbuWraqYTDtuARg/towards-multimodal-interpretability-learning-sparse-2 , accessed 2024-08-12
2024
Closest in time.
Xu, X., Qin, Y., Mi, L., Wang, H., Li, X.: Energy-Based Concept Bottleneck Models: Unifying Prediction, Concept Intervention, and Conditional Interpretations. In: ICLR (2024)
2024
Closest in time.