Fetching the paper…
Reading the bibliography…
In this paper, we review recent approaches for explaining concepts in neural networks.
S. Wermter and W.G. Lehnert, A Hybrid Symbolic/Connectionist Model for Noun Phrase Understanding, Connection Science
1989
Earlier work this paper cites.
S. Wermter, C. Panchev and G. Arevian, Hybrid Neural Plausibility Networks for News Agents, in: Proceedings of the National Conference on Artificial Intelligence AAAI
1999
Earlier work this paper cites.
K.J. McGarry, J. Tait, S. Wermter and J. MacIntyre, Rule-Extraction from Radial Basis Function Networks, in: International Conference on Artificial Neural Networks
1999
Earlier work this paper cites.
S. Wermter, Knowledge Extraction from Transducer Neural Networks, Applied Intelligence
2000
Earlier work this paper cites.
M. Page, Connectionist Modelling in Psychology: A Localist Manifesto, Behavioral and Brain Sciences
2000
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville and Y. Bengio, Generative Adversarial Nets, in: Advances in Neural Information Processing Systems 27
2014
Earlier work this paper cites.
A. Ettinger, A. Elgohary and P. Resnik, Probing for Semantic Evidence of Composition by Means of Simple Classification Tasks, in: Proceedings of the 1st Workshop on Evaluating Vector-Space Representations for NLP
2016
Earlier work this paper cites.
Y. Adi, E. Kermany, Y. Belinkov, O. Lavi and Y. Goldberg, Fine-Grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks, in: International Conference on Learning Representations
2016
Earlier work this paper cites.
D. Bau, B. Zhou, A. Khosla, A. Oliva and A. Torralba, Network Dissection: Quantifying Interpretability of Deep Visual Representations, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
2017
Earlier work this paper cites.
Á. Kádár, G. Chrupała and A. Alishahi, Representation of Linguistic Form and Function in Recurrent Neural Networks, Computational Linguistics
2017
Earlier work this paper cites.
M. Sundararajan, A. Taly and Q. Yan, Axiomatic Attribution for Deep Networks, in: Proceedings of the 34th International Conference on Machine Learning
2017
Earlier work this paper cites.
B. Kim, M. Wattenberg, J. Gilmer, C. Cai, J. Wexler, F. Viegas and R. Sayres, Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV), in: Proceedings of the 35th International Conference on Machine Learning
2018
Earlier work this paper cites.
D. Bau, J.-Y. Zhu, H. Strobelt, B. Zhou, J.B. Tenenbaum, W.T. Freeman and A. Torralba, GAN Dissection: Visualizing and Understanding Generative Adversarial Networks, in: International Conference on Learning Representations
2018
Earlier work this paper cites.
R. Fong and A. Vedaldi, Net2Vec: Quantifying and Explaining How Concepts Are Encoded by Filters in Deep Neural Networks, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
2018
Earlier work this paper cites.
S. Na, Y.J. Choe, D.-H. Lee and G. Kim, Discovery of Natural Language Concepts in Individual Units of CNNs, in: International Conference on Learning Representations
2018
Earlier work this paper cites.
B. Zhou, Y. Sun, D. Bau and A. Torralba, Interpretable Basis Decomposition for Visual Explanation, in: Computer Vision – ECCV 2018
2018
Earlier work this paper cites.
A. Conneau, G. Kruszewski, G. Lample, L. Barrault and M. Baroni, What You Can Cram into a Single \backslash$&!#* Vector: Probing Sentence Embeddings for Linguistic Properties, in: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee and K. Toutanova, BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, in: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei and I. Sutskever, Language Models Are Unsupervised Multitask Learners, 2019. https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf
2019
Earlier work this paper cites.
A. Ghorbani, J. Wexler, J.Y. Zou and B. Kim, Towards Automatic Concept-based Explanations, in: Advances in Neural Information Processing Systems
2019
Earlier work this paper cites.
P.W. Koh, T. Nguyen, Y.S. Tang, S. Mussmann, E. Pierson, B. Kim and P. Liang, Concept Bottleneck Models, in: Proceedings of the 37th International Conference on Machine Learning
2020
Cited alongside, same era.
J. Mu and J. Andreas, Compositional Explanations of Neurons, in: Advances in Neural Information Processing Systems
2020
Cited alongside, same era.
J. Vig, S. Gehrmann, Y. Belinkov, S. Qian, D. Nevo, Y. Singer and S. Shieber, Investigating Gender Bias in Language Models Using Causal Mediation Analysis, in: Advances in Neural Information Processing Systems
2020
Cited alongside, same era.
C.-K. Yeh, B. Kim, S. Arik, C.-L. Li, T. Pfister and P. Ravikumar, On Completeness-aware Concept-Based Explanations in Deep Neural Networks, in: Advances in Neural Information Processing Systems
2020
Cited alongside, same era.
D. Dai, L. Dong, Y. Hao, Z. Sui, B. Chang and F. Wei, Knowledge Neurons in Pretrained Transformers, in: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
2022
Later among the works it cites.
I. Nejadgholi, K. Fraser and S. Kiritchenko, Improving Generalizability in Implicitly Abusive Language Detection with Concept Activation Vectors, in: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
2022
Later among the works it cites.
J. Ferreira, M.d.S. Ribeiro, R. Gonçalves and J. Leite, Looking Inside the Black-Box: Logic-based Explanations for Neural Networks, Proceedings of the International Conference on Principles of Knowledge Representation and Reasoning
2022
Later among the works it cites.
Y. Belinkov, Probing Classifiers: Promises, Shortcomings, and Advances, Computational Linguistics
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Durrani, H. Sajjad, F. Dalvi and Y. Belinkov, Analyzing Individual Neurons in Pre-trained Language Models, in: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)
2020
Cited alongside, same era.
F. Lecue, On the Role of Knowledge Graphs in Explainable AI, Semantic Web
2020
Cited alongside, same era.
2021
Cited alongside, same era.
W. Samek, G. Montavon, S. Lapuschkin, C.J. Anders and K.-R. Müller, Explaining Deep Neural Networks and Beyond: A Review of Methods and Applications, Proceedings of the IEEE
2021
Cited alongside, same era.
A. Radford, J.W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger and I. Sutskever, Learning Transferable Visual Models From Natural Language Supervision, in: Proceedings of the 38th International Conference on Machine Learning
2021
Cited alongside, same era.
M. Finlayson, A. Mueller, S. Gehrmann, S. Shieber, T. Linzen and Y. Belinkov, Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models, in: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)
2021
Cited alongside, same era.
M.d.S. Ribeiro and J. Leite, Aligning Artificial Neural Networks and Ontologies towards Explainable AI, Proceedings of the AAAI Conference on Artificial Intelligence
2021
Cited alongside, same era.
S. Schockaert and V. Gutiérrez-Basulto, Modelling Symbolic Knowledge Using Neural Representations, in: Reasoning Web. Declarative Artificial Intelligence : 17th International Summer School 2021, Leuven, Belgium, September 8–15, 2021, Tutorial Lectures
2022
Cited alongside, same era.
2022
Later among the works it cites.
G. Ciravegna, P. Barbiero, F. Giannini, M. Gori, P. Liò, M. Maggini and S. Melacci, Logic Explained Networks, Artificial Intelligence
2022
Later among the works it cites.
OpenAI, GPT-4 Technical Report, arXiv preprint arXiv:2303.08774
2023
Closest in time.
Y. Jiang, A. Gupta, Z. Zhang, G. Wang, Y. Dou, Y. Chen, L. Fei-Fei, A. Anandkumar, Y. Zhu and L. Fan, VIMA: Robot Manipulation with Multimodal Prompts, in: Proceedings of the 40th International Conference on Machine Learning
2023
Closest in time.
P. Hitzler, M.K. Sarker and A. Eberhart (eds), Compendium of Neurosymbolic Artificial Intelligence
2023
Closest in time.
J.H. Lee, M. Sioutis, K. Ahrens, M. Alirezaie, M. Kerzel and S. Wermter, Chapter 19. Neuro-Symbolic Spatio-Temporal Reasoning, in: Compendium of Neurosymbolic Artificial Intelligence
2023
Closest in time.
R. Dwivedi, D. Dave, H. Naik, S. Singhal, R. Omer, P. Patel, B. Qian, Z. Wen, T. Shah, G. Morgan and R. Ranjan, Explainable AI (XAI): Core Ideas, Techniques, and Solutions, ACM Computing Surveys
2023
Closest in time.
R. Ibrahim and M.O. Shafiq, Explainable Convolutional Neural Networks: A Taxonomy, Review, and Future Directions, ACM Computing Surveys
2023
Closest in time.
F. Sado, C.K. Loo, W.S. Liew, M. Kerzel and S. Wermter, Explainable Goal-driven Agents and Robots - A Comprehensive Review, ACM Computing Surveys
2023
Closest in time.
S. Casper, T. Rauker, A. Ho and D. Hadfield-Menell, Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks, in: First IEEE Conference on Secure and Trustworthy Machine Learning
2023
Closest in time.
T. Oikarinen and T.-W. Weng, CLIP-Dissect: Automatic Description of Neuron Representations in Deep Vision Networks, in: The Eleventh International Conference on Learning Representations
2023
Closest in time.
M. Yuksekgonul, M. Wang and J. Zou, Post-Hoc Concept Bottleneck Models, in: The Eleventh International Conference on Learning Representations
2023
Closest in time.
T. Oikarinen, S. Das, L.M. Nguyen and T.-W. Weng, Label-Free Concept Bottleneck Models, in: The Eleventh International Conference on Learning Representations
2023
Closest in time.
Y. Yang, A. Panagopoulou, S. Zhou, D. Jin, C. Callison-Burch and M. Yatskar, Language in a Bottle: Language Model Guided Concept Bottlenecks for Interpretable Image Classification, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2023
Closest in time.
2023
Closest in time.
P. Barbiero, G. Ciravegna, F. Giannini, M.E. Zarlenga, L.C. Magister, A. Tonda, P. Lio, F. Precioso, M. Jamnik and G. Marra, Interpretable Neural-Symbolic Concept Reasoning, in: Proceedings of the 40th International Conference on Machine Learning
2023
Closest in time.