Fetching the paper…
Reading the bibliography…
Deep learning sequence models have led to a marked increase in performance for a range of Natural Language Processing tasks, but it remains an open question whether they are able to induce proper hierarchical generalizations for representing natural language from linear input alone.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, William W Cohen, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. 2019 · 1901
Earlier work this paper cites.
Neural language models as psycholinguistic subjects: Representations of syntactic state
Richard Futrell, Ethan Wilcox, Takashi Morita, Peng Qian, Miguel Ballesteros, and Roger Levy. 2019 · 1903
Earlier work this paper cites.
Structural supervision improves learning of non-local grammatical dependencies
Ethan Wilcox, Peng Qian, Richard Futrell, Miguel Ballesteros, and Roger Levy. 2019b · 1903
Earlier work this paper cites.
What syntactic structures block dependencies in rnn language models?
Ethan Wilcox, Roger Levy, and Richard Futrell. 2019a · 1905
Earlier work this paper cites.
Three models for the description of language
Noam Chomsky. 1956 · 1956
Earlier work this paper cites.
Constraints on variables in syntax
John Robert Ross. 1967 · 1967
Earlier work this paper cites.
Evidence against the context-freeness of natural language
Stuart M Shieber. 1985 · 1985
Earlier work this paper cites.
Characterizing mildly context-sensitive grammar formalisms
David Jeremy Weir. 1988 · 1988
Earlier work this paper cites.
Finding structure in time
Jeffrey L Elman. 1990 · 1990
Earlier work this paper cites.
Distributed representations, simple recurrent networks, and grammatical structure
Jeffrey L Elman. 1991 · 1991
Earlier work this paper cites.
On multiple context-free grammars
Hiroyuki Seki, Takashi Matsumura, Mamoru Fujii, and Tadao Kasami. 1991 · 1991
Earlier work this paper cites.
100 million words of english: the british national corpus (bnc)
Geoffrey Neil Leech. 1992 · 1992
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Tree-adjoining grammars
Aravind K Joshi and Yves Schabes. 1997 · 1997
Cited alongside, same era.
A probabilistic earley parser as a psycholinguistic model
John Hale. 2001 · 2001
Cited alongside, same era.
Srilm-an extensible language modeling toolkit
Andreas Stolcke. 2002 · 2002
Cited alongside, same era.
Using confidence intervals for graphically based data interpretation
Michael EJ Masson and Geoffrey R Loftus. 2003 · 2003
Cited alongside, same era.
Mixed-effects modeling with crossed random effects for subjects and items
R Harald Baayen, Douglas J Davidson, and Douglas M Bates. 2008 · 2008
Cited alongside, same era.
Expectation-based syntactic comprehension
Roger Levy. 2008 · 2008
Cited alongside, same era.
What do recurrent neural network grammars learn about syntax?
Adhiguna Kuncoro, Miguel Ballesteros, Lingpeng Kong, Chris Dyer, Graham Neubig, and Noah A Smith. 2016 · 2016
Later among the works it cites.
Assessing the ability of lstms to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016 · 2016
Later among the works it cites.
Rnn simulations of grammaticality judgments on long-distance dependencies
Shammur Absar Chowdhury and Roberto Zamparelli. 2018 · 2018
Later among the works it cites.
Rnns as psycholinguistic subjects: Syntactic state and grammatical dependency
Richard Futrell, Ethan Wilcox, Takashi Morita, and Roger Levy. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Random effects structure for confirmatory hypothesis testing: Keep it maximal
Dale J Barr, Roger Levy, Christoph Scheepers, and Harry J Tily. 2013 · 2013
Cited alongside, same era.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson. 2013 · 2013
Cited alongside, same era.
Mildly non-projective dependency grammar
Marco Kuhlmann. 2013 · 2013
Cited alongside, same era.
The effect of word predictability on reading time is logarithmic
Nathaniel J Smith and Roger Levy. 2013 · 2013
Cited alongside, same era.
Parsing as language modeling
Doo Kok Choe and Eugene Charniak. 2016 · 2016
Cited alongside, same era.
Recurrent neural network grammars
Chris Dyer, Adhiguna Kuncoro, Miguel Ballesteros, and Noah A. Smith. 2016 · 2016
Cited alongside, same era.
Mario Giulianelli, Jack Harding, Florian Mohnert, Dieuwke Hupkes, and Willem Zuidema. 2018 · 2018
Later among the works it cites.
Predictive power of word surprisal for reading times is a linear function of language model quality
Adam Goodkind and Klinton Bicknell. 2018 · 2018
Later among the works it cites.
Colorless green recurrent networks dream hierarchically
Kristina Gulordava, Piotr Bojanowski, Edouard Grave, Tal Linzen, and Marco Baroni. 2018 · 2018
Later among the works it cites.
Extracting syntactic trees from transformer encoder self-attentions
David Mareček and Rudolf Rosa. 2018 · 2018
Later among the works it cites.
Targeted syntactic evaluation of language models
Rebecca Marvin and Tal Linzen. 2018 · 2018
Later among the works it cites.
R Thomas McCoy, Robert Frank, and Tal Linzen. 2018 · 2018
Later among the works it cites.
Ordered neurons: Integrating tree structures into recurrent neural networks
Yikang Shen, Shawn Tan, Alessandro Sordoni, and Aaron Courville. 2018 · 2018
Later among the works it cites.
On the practical computational power of finite precision rnns for language recognition
Gail Weiss, Yoav Goldberg, and Eran Yahav. 2018 · 2018
Later among the works it cites.
What do rnn language models learn about filler-gap dependencies?
Ethan Wilcox, Roger Levy, Takashi Morita, and Richard Futrell. 2018 · 2018
Later among the works it cites.