Fetching the paper…
Reading the bibliography…
Analysing whether neural language models encode linguistic information has become popular in NLP.
Assessing BERT’s syntactic abilities
Yoav Goldberg. 2019 · 1901
Earlier work this paper cites.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Syntactic Structures
Noam Chomsky. 1957 · 1957
Earlier work this paper cites.
Shortest connection networks and some generalizations
R. C. Prim. 1957 · 1957
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
358,534 nonwords: The ARC nonword database
Kathleen Rastle, Jonathan Harrington, and Max Coltheart. 2002 · 2002
Earlier work this paper cites.
Non-projective dependency parsing using spanning tree algorithms
Ryan McDonald, Fernando Pereira, Kiril Ribarov, and Jan Hajič. 2005 · 2005
Earlier work this paper cites.
Evaluating the accuracy of an unlexicalized statistical parser on the PARC DepBank
Ted Briscoe and John Carroll. 2006 · 2006
Earlier work this paper cites.
Taming the Jabberwocky: Examining Sentence Processing with Novel Words
Gaurav Kharkwal. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Assessing the ability of LSTMs to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016 · 2016
Earlier work this paper cites.
Universal Dependencies v1: A multilingual treebank collection
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Yoav Goldberg, Jan Hajič, Christopher D. Manning, Ryan McDonald, Slav Petrov, Sampo Pyysalo, Natalia Silveira, Reut Tsarfaty, and Daniel Zeman. 2016 · 2016
Cited alongside, same era.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio. 2017 · 2017
Cited alongside, same era.
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018 · 2018
Cited alongside, same era.
Jabberwocky parsing: Dependency parsing with lexical noise
Jungo Kasai and Robert Frank. 2019 · 2019
Later among the works it cites.
Probing neural network comprehension of natural language arguments
Timothy Niven and Hung-Yu Kao. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Investigating BERT’s knowledge of language: Five analysis methods with NPIs
Alex Warstadt, Yu Cao, Ioana Grosu, Wei Peng, Hagen Blix, Yining Nie, Anna Alsop, Shikha Bordia, Haokun Liu, Alicia Parrish, Sheng-Fu Wang, Jason Phang, Anhad Mohananey, Phu Mon Htut, Paloma Jeretic, and Samuel R. Bowman. 2019 · 2019
Later among the works it cites.
Head-Driven Phrase Structure Grammar parsing on Penn Treebank
Junru Zhou and Hai Zhao. 2019 · 2019
Later among the works it cites.
A tale of a probe and a parser
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kristina Gulordava, Piotr Bojanowski, Edouard Grave, Tal Linzen, and Marco Baroni. 2018 · 2018
Cited alongside, same era.
Visualisation and ‘diagnostic classifiers’ reveal how recurrent and recursive neural networks process hierarchical structure
Dieuwke Hupkes and Willem Zuidema. 2018 · 2018
Cited alongside, same era.
Targeted syntactic evaluation of language models
Rebecca Marvin and Tal Linzen. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D. Manning. 2019 · 2019
Cited alongside, same era.
What does BERT learn about the structure of language?
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. 2019 · 2019
Cited alongside, same era.
Through the Looking-Glass, and What Alice Found There
Lewis Carroll. 1871
Cited in the paper.
Rowan Hall Maudslay, Josef Valvoda, Tiago Pimentel, Adina Williams, and Ryan Cotterell. 2020 · 2020
Later among the works it cites.
Information-theoretic probing for linguistic structure
Tiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod, Adina Williams, and Ryan Cotterell. 2020 · 2020
Later among the works it cites.
A primer in BERTology: What we know about how BERT works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020 · 2020
Later among the works it cites.
Discovering the compositional structure of vector representations with role learning networks
Paul Soulos, R. Thomas McCoy, Tal Linzen, and Paul Smolensky. 2020 · 2020
Later among the works it cites.
SKEP: Sentiment knowledge enhanced pre-training for sentiment analysis
Hao Tian, Can Gao, Xinyan Xiao, Hao Liu, Bolei He, Hua Wu, Haifeng Wang, and Feng Wu. 2020 · 2020
Later among the works it cites.