Fetching the paper…
Reading the bibliography…
Prior work has shown that, on small amounts of training data, syntactic neural language models learn structurally sensitive generalisations more successfully than sequential language models.
Assessing BERT’s syntactic abilities
Yoav Goldberg. 2019 · 1901
Earlier work this paper cites.
Using a stochastic context-free grammar as a language model for speech recognition
D. Jurafsky, C. Wooters, J. Segal, A. Stolcke, E. Fosler, G. Tajchaman, and N. Morgan. 1995 · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Statistical Methods for Speech Recognition
Frederick Jelinek. 1997 · 1997
Earlier work this paper cites.
Supertagging: An approach to almost parsing
Srinivas Bangalore and Aravind K. Joshi. 1999 · 1999
Earlier work this paper cites.
Structured language modeling
Ciprian Chelba and Frederick Jelinek. 2000 · 2000
Earlier work this paper cites.
Two decades of statistical language modeling: Where do we go from here
Ronald Rosenfeld. 2000 · 2000
Earlier work this paper cites.
Probabilistic top-down parsing and language modeling
Brian Roark. 2001 · 2001
Earlier work this paper cites.
Discriminative training of a neural network statistical parser
James Henderson. 2004 · 2004
Earlier work this paper cites.
A neural syntactic language model
Ahmad Emami and Frederick Jelinek. 2005 · 2005
Earlier work this paper cites.
Model compression
Cristian Bucilǎ, Rich Caruana, and Alexandru Niculescu-Mizil. 2006 · 2006
Earlier work this paper cites.
Wide-coverage efficient statistical parsing with CCG and log-linear models
Stephen Clark and James R. Curran. 2007 · 2007
Earlier work this paper cites.
Improved inference for unlexicalized parsing
Slav Petrov and Dan Klein. 2007 · 2007
Earlier work this paper cites.
Statistical Machine Translation
Philipp Koehn. 2010 · 2010
Earlier work this paper cites.
Recurrent neural network based language model
Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černocký, and Sanjeev Khudanpur. 2010 · 2010
Earlier work this paper cites.
Do deep nets really need to be deep?
Jimmy Ba and Rich Caruana. 2014 · 2014
Earlier work this paper cites.
A Bayesian model for generative transition-based dependency parsing
Jan Buys and Phil Blunsom. 2015 · 2015
Earlier work this paper cites.
Transition-based dependency parsing with stack long short-term memory
Chris Dyer, Miguel Ballesteros, Wang Ling, Austin Matthews, and Noah A. Smith. 2015 · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 · 2015
Earlier work this paper cites.
Dependency recurrent neural language models for sentence completion
Piotr Mirowski and Andreas Vlachos. 2015 · 2015
Earlier work this paper cites.
Parsing as language modeling
Do Kook Choe and Eugene Charniak. 2016 · 2016
Cited alongside, same era.
Recurrent neural network grammars
Chris Dyer, Adhiguna Kuncoro, Miguel Ballesteros, and Noah A. Smith. 2016 · 2016
Cited alongside, same era.
Tree-to-sequence attentional neural machine translation
Akiko Eriguchi, Kazuma Hashimoto, and Yoshimasa Tsuruoka. 2016 · 2016
Cited alongside, same era.
Sequence-level knowledge distillation
Yoon Kim and Alexander M. Rush. 2016 · 2016
Cited alongside, same era.
Assessing the ability of LSTMs to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016 · 2016
Cited alongside, same era.
Multi-task sequence to sequence learning
Thang Luong, Quoc V. Le, Ilya Sutskever, Oriol Vinyals, and Lukasz Kaiser. 2016 · 2016
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
Deep rnns encode soft hierarchical syntax
Terra Blevins, Omer Levy, and Luke Zettlemoyer and. 2018 · 2018
Later among the works it cites.
What you can cram into a single vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, Germán Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018 · 2018
Later among the works it cites.
Born-again neural networks
Tommaso Furlanello, Zachary Chase Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar. 2018 · 2018
Later among the works it cites.
Colorless green recurrent networks dream hierarchically
Kristina Gulordava, Piotr Bojanowski, Edouard Grave, Tal Linzen, and Marco Baroni. 2018 · 2018
Later among the works it cites.
Finding syntax in human encephalography with beam search
John Hale, Chris Dyer, Adhiguna Kuncoro, and Jonathan R. Brennan. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Does string-based neural MT learn source syntax?
Xing Shi, Inkit Padhi, and Kevin Knight. 2016 · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna. 2016 · 2016
Cited alongside, same era.
Fine-grained analysis of sentence embeddings using auxiliary prediction tasks
Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg. 2017 · 2017
Cited alongside, same era.
Towards string-to-tree neural machine translation
Roee Aharoni and Yoav Goldberg. 2017 · 2017
Cited alongside, same era.
What do neural machine translation models learn about morphology?
Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, and James Glass. 2017 · 2017
Cited alongside, same era.
Hierarchical multiscale recurrent neural networks
Junyoung Chung, Sungjin Ahn, and Yoshua Bengio. 2017 · 2017
Cited alongside, same era.
Later among the works it cites.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder. 2018 · 2018
Later among the works it cites.
Lstms can learn syntax-sensitive dependencies well, but modeling structure makes them better
Adhiguna Kuncoro, Chris Dyer, John Hale, Dani Yogatama, Stephen Clark, and Phil Blunsom. 2018 · 2018
Later among the works it cites.
Distilling knowledge for search-based structured prediction
Yijia Liu, Wanxiang Che, Huaipeng Zhao, Bing Qin, and Ting Liu. 2018 · 2018
Later among the works it cites.
Targeted syntactic evaluation of language models
Rebecca Marvin and Tal Linzen. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
The importance of being recurrent for modeling hierarchical structure
Ke M. Tran, Arianna Bisazza, and Christof Monz. 2018 · 2018
Later among the works it cites.
Memory architectures in recurrent neural network language models
Dani Yogatama, Yishu Miao, Gabor Melis, Wang Ling, Adhiguna Kuncoro, Chris Dyer, and Phil Blunsom. 2018 · 2018
Later among the works it cites.
Universal transformers
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser. 2019 · 2019
Closest in time.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Closest in time.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D. Manning. 2019 · 2019
Closest in time.
Unsupervised recurrent neural network grammars
Yoon Kim, Alexander M. Rush, Lei Yu, Adhiguna Kuncoro, Chris Dyer, and Gabor Melis. 2019 · 2019
Closest in time.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Closest in time.
Ordered neurons: Integrating tree structures into recurrent neural networks
Yikang Shen, Shawn Tan, Alessandro Sordoni, and Aaron C. Courville. 2019 · 2019
Closest in time.