Fetching the paper…
Reading the bibliography…
Natural language exhibits patterns of hierarchically governed dependencies, in which relations between words are sensitive to syntactic structure rather than linear ordering.
Assessing BERT’s syntactic abilities
Yoav Goldberg. 2019 · 1901
Earlier work this paper cites.
Problems of Knowledge and Freedom
Noam Chomsky. 1971 · 1971
Earlier work this paper cites.
Is structure dependence an innate constraint? new experimental evidence from children’s complex-question production
Ben Ambridge, Caroline F. Rowland, and Julian M. Pine. 2008 · 2008
Earlier work this paper cites.
The learnability of abstract syntactic principles
Amy Perfors, Joshua B Tenenbaum, and Terry Regier. 2011 · 2011
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Minh-Thang Luong, Hieu Pham, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2016 · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, undefinedukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Colorless green recurrent networks dream hierarchically
Kristina Gulordava, Piotr Bojanowski, Edouard Grave, Tal Linzen, and Marco Baroni. 2018 · 2018
Earlier work this paper cites.
LSTMs can learn syntax-sensitive dependencies well, but modeling structure makes them better
Adhiguna Kuncoro, Chris Dyer, John Hale, Dani Yogatama, Stephen Clark, and Phil Blunsom. 2018 · 2018
Earlier work this paper cites.
Targeted syntactic evaluation of language models
Rebecca Marvin and Tal Linzen. 2018 · 2018
Cited alongside, same era.
Ordered neurons: Integrating tree structures into recurrent neural networks
Yikang Shen, Shawn Tan, Alessandro Sordoni, and Aaron Courville. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D. Manning. 2019 · 2019
Cited alongside, same era.
Unsupervised recurrent neural network grammars
Yoon Kim, Alexander Rush, Lei Yu, Adhiguna Kuncoro, Chris Dyer, and Gábor Melis. 2019 · 2019
Cited alongside, same era.
Investigating BERT’s knowledge of language: Five analysis methods with npis
Alex Warstadt, Yu Cao, Ioana Grosu, Wei Peng, Hagen Blix, Yining Nie, Anna Alsop, Shikha Bordia, Haokun Liu, Alicia Parrish, et al. 2019 · 2019
Later among the works it cites.
Sequence-to-sequence networks learn the meaning of reflexive anaphora
Robert Frank and Jackson Petty. 2020 · 2020
Later among the works it cites.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn. 2020 · 2020
Later among the works it cites.
A systematic assessment of syntactic generalization in neural language models
Jennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox, and Roger Levy. 2020 · 2020
Later among the works it cites.
COGS: A compositional generalization challenge based on semantic interpretation
Najoung Kim and Tal Linzen. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Open sesame: Getting inside BERT’s linguistic knowledge
Yongjie Lin, Yi Chern Tan, and Robert Frank. 2019 · 2019
Cited alongside, same era.
Sequential neural networks as automata
William Merrill. 2019 · 2019
Cited alongside, same era.
Finding hierarchical structure in neural stacks using unsupervised parsing
William Merrill, Lenny Khazan, Noah Amsel, Yiding Hao, Simon Mendelsohn, and Robert Frank. 2019 · 2019
Cited alongside, same era.
Quantity doesn’t buy quality syntax with neural language models
Marten Van Schijndel, Aaron Mueller, and Tal Linzen. 2019 · 2019
Cited alongside, same era.
R. Thomas McCoy, Robert Frank, and Tal Linzen. 2020 · 2020
Later among the works it cites.
Can neural networks acquire a structural bias from raw linguistic data?
Alex Warstadt and Samuel R. Bowman. 2020 · 2020
Later among the works it cites.
A primer in bertology: What we know about how bert works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2021 · 2021
Closest in time.