Fetching the paper…
Reading the bibliography…
The proliferation of deep neural networks in various domains has seen an increased need for interpretability of these models.
Discovery of natural language concepts in individual units of CNNs
Seil Na, Yo Joong Choe, Dong-Hyun Lee, and Gunhee Kim. 2019 · 1902
Earlier work this paper cites.
Unsupervised transfer learning via BERT neuron selection
Mehrdad Valipour, En-Shiun Annie Lee, Jaime R. Jamacaro, and Carolina Bessega. 2019 · 1912
Earlier work this paper cites.
A survey of formal grammars and algorithms for recognition and transformation in mechanical translation
Bernard Vauquois. 1968 · 1968
Earlier work this paper cites.
Concepts and semantic relations in information science
Wolfgang G Stock. 2010 · 1969
Earlier work this paper cites.
Richard Meyes, Constantin Waubert de Puiseau, Andres Posada-Moreno, and Tobias Meisen. 2020 · 2004
Earlier work this paper cites.
Finding experts in transformer models
Xavier Suau, Luca Zappella, and Nicholas Apostoloff. 2020 · 2005
Earlier work this paper cites.
Compositional explanations of neurons
Jesse Mu and Jacob Andreas. 2020 · 2006
Earlier work this paper cites.
Computing optimal subsets
Maxim Binshtok, Ronen I Brafman, Solomon Eyal Shimony, Ajay Martin, and Crag Boutilier. 2007 · 2007
Earlier work this paper cites.
Visualizing higher-layer features of a deep network
Dumitru Erhan, Yoshua Bengio, Aaron Courville, and Pascal Vincent. 2009 · 2009
Earlier work this paper cites.
Unsupervised feature learning and deep learning: A review and new perspectives
Yoshua Bengio, Aaron C. Courville, and Pascal Vincent. 2012 · 2012
Earlier work this paper cites.
Semantics-based machine translation with hyperedge replacement grammars
Bevan Jones, Jacob Andreas, Daniel Bauer, Karl Moritz Hermann, and Kevin Knight. 2012 · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Neural Machine Translation by Jointly Learning to Align and Translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014 · 2014
Earlier work this paper cites.
Sparse overcomplete word vector representations
Manaal Faruqui, Yulia Tsvetkov, Dani Yogatama, Chris Dyer, and Noah A. Smith. 2015 · 2015
Earlier work this paper cites.
A compositional and interpretable semantic space
Alona Fyshe, Leila Wehbe, Partha P. Talukdar, Brian Murphy, and Tom M. Mitchell. 2015 · 2015
Earlier work this paper cites.
Visualizing and understanding recurrent networks
Andrej Karpathy, Justin Johnson, and Li Fei-Fei. 2015 · 2015
Earlier work this paper cites.
Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks
Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg. 2016 · 2016
Earlier work this paper cites.
Visualizing and understanding neural models in NLP
Jiwei Li, Xinlei Chen, Eduard Hovy, and Dan Jurafsky. 2016a · 2016
Earlier work this paper cites.
The Mythos of Model Interpretability
Zachary C Lipton. 2016 · 2016
Earlier work this paper cites.
Representation of linguistic form and function in recurrent neural networks
Ákos Kádár, Grzegorz Chrupała, and Afra Alishahi. 2017 · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee. 2017 · 2017
Earlier work this paper cites.
SVCCA: Singular Vector Canonical Correlation Analysis for Deep Learning Dynamics and Interpretability
Maithra Raghu, Justin Gilmer, Jason Yosinski, and Jascha Sohl-Dickstein. 2017 · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Earlier work this paper cites.
Visualizing deep neural network decisions: Prediction difference analysis
Luisa M. Zintgraf, Taco S. Cohen, Tameem Adel, and Max Welling. 2017 · 2017
Earlier work this paper cites.
Understanding neural networks and individual neuron importance via information-ordered cumulative ablation
Rana Ali Amjad, Kairen Liu, and Bernhard C. Geiger. 2018 · 2018
Cited alongside, same era.
What you can cram into a single vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018 · 2018
Cited alongside, same era.
Kedar Dhamdhere, Mukund Sundararajan, and Qiqi Yan. 2018 · 2018
Cited alongside, same era.
Pathologies of neural models make interpretations difficult
Shi Feng, Eric Wallace, Alvin Grissom II, Mohit Iyyer, Pedro Rodriguez, and Jordan Boyd-Graber. 2018 · 2018
Cited alongside, same era.
Explaining character-aware neural networks for word-level prediction: Do they discover linguistic rules?
Fréderic Godin, Kris Demuynck, Joni Dambre, Wesley De Neve, and Thomas Demeester. 2018 · 2018
Cited alongside, same era.
How do decisions emerge across layers in neural models? interpretation with differentiable masking
Nicola De Cao, Michael Sejr Schlichtkrull, Wilker Aziz, and Ivan Titov. 2020 · 2020
Later among the works it cites.
The shapley taylor interaction index
Kedar Dhamdhere, Ashish Agarwal, and Mukund Sundararajan. 2020 · 2020
Later among the works it cites.
Analyzing individual neurons in pre-trained language models
Nadir Durrani, Hassan Sajjad, Fahim Dalvi, and Yonatan Belinkov. 2020 · 2020
Later among the works it cites.
Intrinsic probing through dimension selection
Lucas Torroba Hennigen, Adina Williams, and Ryan Cotterell. 2020 · 2020
Later among the works it cites.
Hierarchy and interpretability in neural models of language processing
Dieuwke Hupkes. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dieuwke Hupkes, Sara Veldhoen, and Willem Zuidema. 2018 · 2018
Cited alongside, same era.
Nlp’s generalization problem, and how researchers are tackling it
Ana Marasović. 2018 · 2018
Cited alongside, same era.
Beyond word importance: Contextual decomposition to extract interactions from lstms
W. James Murdoch, Peter J. Liu, and Bin Yu. 2018 · 2018
Cited alongside, same era.
The building blocks of interpretability
Chris Olah, Arvind Satyanarayan, Ian Johnson, Shan Carter, Ludwig Schubert, Katherine Ye, and Alexander Mordvintsev. 2018 · 2018
Cited alongside, same era.
Interpretable textual neuron representations for NLP
Nina Poerner, Benjamin Roth, and Hinrich Schütze. 2018 · 2018
Cited alongside, same era.
The Importance of Being Recurrent for Modeling Hierarchical Structure
Ke Tran, Arianna Bisazza, and Christof Monz. 2018 · 2018
Cited alongside, same era.
Language modeling teaches you more than translation does: Lessons learned through auxiliary syntactic task analysis
Kelly Zhang and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
Leland Mclnnes, John Healy, and James Melville. 2020 · 2020
Later among the works it cites.
Asking without telling: Exploring latent ontologies in contextual representations
Julian Michael, Jan A. Botha, and Ian Tenney. 2020 · 2020
Later among the works it cites.
Information-theoretic probing for linguistic structure
Tiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod, Adina Williams, and Ryan Cotterell. 2020 · 2020
Later among the works it cites.
Information-theoretic probing with minimum description length
Elena Voita and Ivan Titov. 2020 · 2020
Later among the works it cites.
Similarity Analysis of Contextual Word Representation Models
John Wu, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2020 · 2020
Later among the works it cites.
Ecco: An open source library for the explainability of transformer language models
J Alammar. 2021 · 2021
Closest in time.
Knowledge neurons in pretrained transformers
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, and Furu Wei. 2021 · 2021
Closest in time.
How transfer learning impacts linguistic knowledge in deep nlp models?
Nadir Durrani, Hassan Sajjad, and Fahim Dalvi. 2021 · 2021
Closest in time.
Amnesic probing: Behavioral explanation with amnesic counterfactuals
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg. 2021 · 2021
Closest in time.
CausaLM: Causal model explanation through counterfactual language models
Amir Feder, Nadav Oved, Uri Shalit, and Roi Reichart. 2021 · 2021
Closest in time.
Pruning-then-expanding model for domain adaptation of neural machine translation
Shuhao Gu, Yang Feng, and Wanying Xie. 2021 · 2021
Closest in time.
Effect of post-processing on contextualized word representations
Hassan Sajjad, Firoj Alam, Fahim Dalvi, and Nadir Durrani. 2021 · 2021
Closest in time.
Implicit representations of event properties within contextual language models: Searching for “causativity neurons"
Esther Seyffarth, Younes Samih, Laura Kallmeyer, and Hassan Sajjad. 2021 · 2021
Closest in time.
On the pitfalls of analyzing individual neurons in language models
Omer Antverg and Yonatan Belinkov. 2022 · 2022
Closest in time.
Idani: Inference-time domain adaptation via neuron-level interventions
Omer Antverg, Eyal Ben-David, and Yonatan Belinkov. 2022 · 2022
Closest in time.
Discovering latent concepts learned in BERT
Fahim Dalvi, Abdul Rafae Khan, Firoj Alam, Nadir Durrani, Jia Xu, and Hassan Sajjad. 2022 · 2022
Closest in time.
Linguistic correlation analysis: Discovering salient neurons in deepnlp models
Nadir Durrani, Fahim Dalvi, and Hassan Sajjad. 2022 · 2022
Closest in time.
Analyzing encoded concepts in transformer language models
Hassan Sajjad, Nadir Durrani, Fahim Dalvi, Firoj Alam, Abdul Khan, and Jia Xu. 2022 · 2022
Closest in time.
A latent-variable model for intrinsic probing
Karolina Stanczak, Lucas Torroba Hennigen, Adina Williams, Ryan Cotterell, and Isabelle Augenstein. 2022 · 2022
Closest in time.