Fetching the paper…
Reading the bibliography…
We describe a procedure for explaining neurons in deep representations by identifying compositional logical concepts that closely approximate neuron behavior.
Inductive programming: A survey of program synthesis techniques
E. Kitzelmann · 2009
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
K. Simonyan, A. Vedaldi, and A. Zisserman · 2013
Earlier work this paper cites.
GloVe: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Earlier work this paper cites.
Intriguing properties of neural networks
C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
M. D. Zeiler and R. Fergus · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
S. Bowman, G. Angeli, C. Potts, and C. D. Manning · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Object detectors emerge in deep scene CNNs
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba · 2015
Earlier work this paper cites.
A fast unified model for parsing and sentence understanding
S. Bowman, J. Gauthier, A. Rastogi, R. Gupta, C. D. Manning, and C. Potts · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Generating visual explanations
L. A. Hendricks, Z. Akata, M. Rohrbach, J. Donahue, B. Schiele, and T. Darrell · 2016
Earlier work this paper cites.
Synthesizing the preferred inputs for neurons in neural networks via deep generator networks
A. Nguyen, A. Dosovitskiy, J. Yosinski, T. Brox, and J. Clune · 2016
Cited alongside, same era.
Translating neuralese
J. Andreas, A. Dragan, and D. Klein · 2017
Cited alongside, same era.
Network dissection: Quantifying interpretability of deep visual representations
D. Bau, B. Zhou, A. Khosla, A. Oliva, and A. Torralba · 2017
Cited alongside, same era.
Densely connected convolutional networks
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger · 2017
Cited alongside, same era.
Generating natural language adversarial examples
M. Alzantot, Y. S. Sharma, A. Elgohary, B.-J. Ho, M. Srivastava, and K.-W. Chang · 2018
Cited alongside, same era.
e-SNLI: natural language inference with natural language explanations
O.-M. Camburu, T. Rocktäschel, T. Lukasiewicz, and P. Blunsom · 2018
Activation atlas
S. Carter, Z. Armstrong, L. Schubert, I. Johnson, and C. Olah · 2019
Later among the works it cites.
Designing and interpreting probes with control tasks
J. Hewitt and P. Liang · 2019
Later among the works it cites.
A structural probe for finding syntax in word representations
J. Hewitt and C. D. Manning · 2019
Later among the works it cites.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
T. McCoy, E. Pavlick, and T. Linzen · 2019
Later among the works it cites.
Explain yourself! Leveraging language models for commonsense reasoning
N. F. Rajani, B. McCann, C. Xiong, and R. Socher · 2019
Later among the works it cites.
Do ImageNet classifiers generalize to ImageNet?
B. Recht, R. Roelofs, L. Schmidt, and V. Shankar · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Net2vec: Quantifying and explaining how concepts are encoded by filters in deep neural networks
R. Fong and A. Vedaldi · 2018
Cited alongside, same era.
Annotation artifacts in natural language inference data
S. Gururangan, S. Swayamdipta, O. Levy, R. Schwartz, S. Bowman, and N. A. Smith · 2018
Cited alongside, same era.
Grounding visual explanations
L. A. Hendricks, R. Hu, T. Darrell, and Z. Akata · 2018
Cited alongside, same era.
Hypothesis only baselines in natural language inference
A. Poliak, J. Naradowsky, A. Haldar, R. Rudinger, and B. Van Durme · 2018
Cited alongside, same era.
ObjectNet: A large-scale bias-controlled dataset for pushing the limits of object recognition models
A. Barbu, D. Mayo, J. Alverio, W. Luo, C. Wang, D. Gutfreund, J. Tenenbaum, and B. Katz · 2019
Cited alongside, same era.
Identifying and controlling important neurons in neural machine translation
A. Bau, Y. Belinkov, H. Sajjad, N. Durrani, F. Dalvi, and J. Glass · 2019
Cited alongside, same era.
Universal adversarial triggers for attacking and analyzing nlp
E. Wallace, S. Feng, N. Kandpal, M. Gardner, and S. Singh · 2019
Later among the works it cites.
Evaluating NLP models via contrast sets
M. Gardner, Y. Artzi, V. Basmova, J. Berant, B. Bogin, S. Chen, P. Dasigi, D. Dua, Y. Elazar, A. Gottumukkala, et al · 2020
Closest in time.
Learning the difference that makes a difference with counterfactually-augmented data
D. Kaushik, E. Hovy, and Z. C. Lipton · 2020
Closest in time.
Zoom in: An introduction to circuits
C. Olah, N. Cammarata, L. Schubert, G. Goh, M. Petrov, and S. Carter · 2020
Closest in time.
Information-theoretic probing for linguistic structure
T. Pimentel, J. Valvoda, R. H. Maudslay, R. Zmigrod, A. Williams, and R. Cotterell · 2020
Closest in time.