Fetching the paper…
Reading the bibliography…
While a lot of analysis has been carried to demonstrate linguistic knowledge captured by the representations learned within deep NLP models, very little attention has been paid towards individual neurons.We carry outa neuron-level analysis using core linguistic tasks of predicting morphology, syntax and semantics, on pre-trained language models, with questions like: i) do individual neurons in pre-trained models capture linguistic information? ii) which parts of the network learn more about certain linguistic phenomena? iii) how distributed or focused is the information? and iv) how do various architectures differ in learning these properties? We found small subsets of neurons to predict linguistic tasks, with lower level tasks (such as morphology) localized in fewer neurons, compared to higher level task of predicting syntax.
Building a large annotated corpus of English: The Penn Treebank
Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz. 1993 · 1993
Earlier work this paper cites.
Supertagging: An approach to almost parsing
Srinivas Bangalore and Aravind K. Joshi. 1999 · 1999
Earlier work this paper cites.
Introduction to the CoNLL-2000 shared task chunking
Erik F. Tjong Kim Sang and Sabine Buchholz. 2000 · 2000
Earlier work this paper cites.
Finding experts in transformer models
Xavier Suau, Luca Zappella, and Nicholas Apostoloff. 2020 · 2005
Earlier work this paper cites.
Regularization and variable selection via the elastic net
Hui Zou and Trevor Hastie. 2005 · 2005
Earlier work this paper cites.
Creating a CCGbank and a wide-coverage CCG lexicon for German
Julia Hockenmaier. 2006 · 2006
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Visualizing and understanding recurrent networks
Andrej Karpathy, Justin Johnson, and Li Fei-Fei. 2015 · 2015
Earlier work this paper cites.
Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks
Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg. 2016 · 2016
Earlier work this paper cites.
Visualizing and understanding neural models in NLP
Jiwei Li, Xinlei Chen, Eduard Hovy, and Dan Jurafsky. 2016 · 2016
Earlier work this paper cites.
Analyzing Linguistic Knowledge in Sequential Model of Sentence
Peng Qian, Xipeng Qiu, and Xuanjing Huang. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Does string-based neural MT learn source syntax?
Xing Shi, Inkit Padhi, and Kevin Knight. 2016 · 2016
Earlier work this paper cites.
The parallel meaning bank: Towards a multilingual corpus of translations annotated with compositional meaning representations
Lasha Abzianidze, Johannes Bjerva, Kilian Evang, Hessel Haagsma, Rik van Noord, Pierre Ludmann, Duc-Duy Nguyen, and Johan Bos. 2017 · 2017
Earlier work this paper cites.
Understanding and Improving Morphological Learning in the Neural Machine Translation Decoder
Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, and Stephan Vogel. 2017 · 2017
Cited alongside, same era.
Representation of linguistic form and function in recurrent neural networks
Akos Kádár, Grzegorz Chrupała, and Afra Alishahi. 2017 · 2017
Cited alongside, same era.
Deep RNNs encode soft hierarchical syntax
Terra Blevins, Omer Levy, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
What you can cram into a single vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018 · 2018
Cited alongside, same era.
Filter-wrapper combination and embedded feature selection for gene expression data
Shilan Hameed. 2018 · 2018
Cited alongside, same era.
Deep contextualized word representations
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang. 2019 · 2019
Later among the works it cites.
Linguistic knowledge and transferability of contextual representations
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith. 2019 · 2019
Later among the works it cites.
To tune or not to tune? adapting pretrained representations to diverse tasks
Matthew E. Peters, Sebastian Ruder, and Noah A. Smith. 2019 · 2019
Later among the works it cites.
What do you learn from context? probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Sam Bowman, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Later among the works it cites.
A multiscale visualization of attention in the transformer model
Jesse Vig. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018a · 2018
Cited alongside, same era.
Dissecting contextual word embeddings: Architecture and representation
Matthew Peters, Mark Neumann, Luke Zettlemoyer, and Wen-tau Yih. 2018b · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
Identifying and controlling important neurons in neural machine translation
Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2019 · 2019
Cited alongside, same era.
Analysis methods in neural language processing: A survey
Yonatan Belinkov and James Glass. 2019 · 2019
Cited alongside, same era.
What does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Cited alongside, same era.
What is one grain of sand in the desert? analyzing individual neurons in deep nlp models
Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, D. Anthony Bau, and James Glass. 2019 · 2019
Cited alongside, same era.
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 2019
Later among the works it cites.
On the linguistic representational power of neural machine translation models
Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, and James Glass. 2020 · 2020
Closest in time.
Analyzing redundancy in pretrained transformer models
Fahim Dalvi, Hassan Sajjad, Nadir Durrani, and Yonatan Belinkov. 2020 · 2020
Closest in time.
Are pre-trained language models aware of phrases? simple but strong baselines for grammar induction
Taeuk Kim, Jihun Choi, Daniel Edmiston, and Sang goo Lee. 2020 · 2020
Closest in time.
Information-theoretic probing for linguistic structure
Tiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod, Adina Williams, and Ryan Cotterell. 2020 · 2020
Closest in time.
Tx-ray: Quantifying and explaining model-knowledge transfer in (un-)supervised NLP
Nils Rethmeier, Vageesh Kumar Saxena, and Isabelle Augenstein. 2020 · 2020
Closest in time.
Information-theoretic probing with minimum description length
Elena Voita and Ivan Titov. 2020 · 2020
Closest in time.
Similarity Analysis of Contextual Word Representation Models
John Wu, Hassan Belinkov, Yonatan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2020 · 2020
Closest in time.