Fetching the paper…
Reading the bibliography…
Faithfully summarizing the knowledge encoded by a deep neural network (DNN) into a few symbolic primitive patterns without losing much information represents a core challenge in explainable AI.
A simplified bargaining model for the n-person cooperative game
John C Harsanyi · 1963
Earlier work this paper cites.
The mnist database of handwritten digits
Yann LeCun · 1998
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Explaining generalization power of a dnn using interactive concepts
Huilin Zhou, Hao Zhang, Huiqi Deng, Dongrui Liu, Wen Shen, Shih-Han Chan, and Quanshi Zhang · 2013
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Understanding neural networks through deep visualization
Jason Yosinski, Jeff Clune, Anh Nguyen, Thomas Fuchs, and Hod Lipson · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Cited alongside, same era.
Network dissection: Quantifying interpretability of deep visual representations
David Bau, Bolei Zhou, Aditya Khosla, Aude Oliva, and Antonio Torralba · 2017
Cited alongside, same era.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra · 2017
Cited alongside, same era.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Rosetta neurons: Mining the common units in a model zoo
Amil Dravid, Yossi Gandelsman, Alexei A. Efros, and Assaf Shocher · 2023
Later among the works it cites.
Can the inference logic of large language models be disentangled into symbolic concepts?
Wen Shen, Lei Cheng, Yuxiao Yang, Mingjie Li, and Quanshi Zhang · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Explaining how a neural network play the go game and let people learn
Huilin Zhou, Huijie Tang, Mingjie Li, Hao Zhang, Zhenyu Liu, and Quanshi Zhang · 2023
Later among the works it cites.
Unifying fourteen post-hoc attribution methods with taylor interactions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Discovering and explaining the representation bottleneck of dnns
Huiqi Deng, Qihan Ren, Hao Zhang, and Quanshi Zhang · 2022
Cited alongside, same era.
Aquila-7b
BAAI · 2023
Cited alongside, same era.
Mingjie Li and Quanshi Zhang
Cited in the paper.
Does a neural network really encode symbolic concept?
Mingjie Li and Quanshi Zhang
Cited in the paper.
Defining and quantifying the emergence of sparse concepts in dnns
Jie Ren, Mingjie Li, Qirui Chen, Huiqi Deng, and Quanshi Zhang
Cited in the paper.
Can we faithfully represent masked states to compute shapley values on a dnn?
Jie Ren, Zhanpeng Zhou, Qirui Chen, and Quanshi Zhang
Cited in the paper.
Bayesian neural networks avoid encoding perturbation-sensitive and complex concepts
Qihan Ren, Huiqi Deng, Yunuo Chen, Siyu Lou, and Quanshi Zhang
Cited in the paper.
Huiqi Deng, Na Zou, Mengnan Du, Weifu Chen, Guocan Feng, Ziwei Yang, Zheyang Li, and Quanshi Zhang · 2024
Closest in time.
Towards the difficulty for a deep neural network to learn concepts of different complexities
Dongrui Liu, Huiqi Deng, Xu Cheng, Qihan Ren, Kangrui Wang, and Quanshi Zhang · 2024
Closest in time.
Where we have arrived in proving the emergence of sparse interaction primitives in dnns
Qihan Ren, Jiayang Gao, Wen Shen, and Quanshi Zhang · 2024
Closest in time.