Fetching the paper…
Reading the bibliography…
Recursion is a prominent feature of human language, and fundamentally challenging for self-attention due to the lack of an explicit recursive-state tracking mechanism.
Memory-augmented recurrent neural networks can learn generalized dyck languages
Mirac Suzgun, Sebastian Gehrmann, Yonatan Belinkov, and Stuart M Shieber. 2019 · 1911
Earlier work this paper cites.
Bllip 1987-89 wsj corpus release 1
Eugene Charniak, Don Blaheta, Niyu Ge, Keith Hall, John Hale, and Mark Johnson. 2000 · 1987
Earlier work this paper cites.
Learning context-free grammars: Capabilities and limitations of a recurrent neural network with an external stack memory
Sreerupa Das, C Lee Giles, and Guo-Zheng Sun. 1992 · 1992
Earlier work this paper cites.
A structured language model
Ciprian Chelba. 1997 · 1997
Earlier work this paper cites.
Three generative, lexicalised models for statistical parsing
Michael Collins. 1997 · 1997
Earlier work this paper cites.
The faculty of language: What is it, who has it, and how did it evolve?
Marc D. Hauser, Noam Chomsky, and W. Tecumseh Fitch. 2002 · 2002
Earlier work this paper cites.
Learning to transduce with unbounded memory
Edward Grefenstette, Karl Moritz Hermann, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Earlier work this paper cites.
Inferring algorithmic patterns with stack-augmented recurrent nets
Armand Joulin and Tomas Mikolov. 2015 · 2015
Earlier work this paper cites.
Karol Kurach, Marcin Andrychowicz, and Ilya Sutskever. 2015 · 2015
Earlier work this paper cites.
Dependency recurrent neural language models for sentence completion
Piotr Mirowski and Andreas Vlachos. 2015 · 2015
Earlier work this paper cites.
Grammar as a foreign language
Oriol Vinyals, Łukasz Kaiser, Terry Koo, Slav Petrov, Ilya Sutskever, and Geoffrey Hinton. 2015 · 2015
Earlier work this paper cites.
A fast unified model for parsing and sentence understanding
Samuel Bowman, Jon Gauthier, Abhinav Rastogi, Raghav Gupta, Christopher D Manning, and Christopher Potts. 2016 · 2016
Earlier work this paper cites.
Parsing as language modeling
Do Kook Choe and Eugene Charniak. 2016 · 2016
Earlier work this paper cites.
Recurrent neural network grammars
Chris Dyer, Adhiguna Kuncoro, Miguel Ballesteros, and Noah A. Smith. 2016 · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017 · 2017
Earlier work this paper cites.
Effective inference for generative neural parsing
Mitchell Stern, Daniel Fried, and Dan Klein. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Targeted syntactic evaluation of language models
Rebecca Marvin and Tal Linzen. 2018 · 2018
Cited alongside, same era.
Linguistically-informed self-attention for semantic role labeling
Emma Strubell, Patrick Verga, Daniel Andor, David Weiss, and Andrew McCallum. 2018 · 2018
Cited alongside, same era.
What does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Cited alongside, same era.
Unsupervised latent tree induction with deep inside-outside recursive auto-encoders
Andrew Drozdov, Patrick Verga, Mohit Yadav, Mohit Iyyer, and Andrew McCallum. 2019 · 2019
Theoretical Limitations of Self-Attention in Neural Sequence Models
Michael Hahn. 2020 · 2020
Later among the works it cites.
A systematic assessment of syntactic generalization in neural language models
Jennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox, and Roger Levy. 2020 · 2020
Later among the works it cites.
Emergent linguistic structure in artificial neural networks trained by self-supervision
Christopher D. Manning, Kevin Clark, John Hewitt, Urvashi Khandelwal, and Omer Levy. 2020 · 2020
Later among the works it cites.
BLiMP: The benchmark of linguistic minimal pairs for English
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R. Bowman. 2020 · 2020
Later among the works it cites.
Attention is turing-complete
Jorge Perez, Pablo Barcelo, and Javier Marinkovic. 2021 · 2021
Later among the works it cites.
Structural guidance for transformer language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D. Manning. 2019 · 2019
Cited alongside, same era.
Unsupervised recurrent neural network grammars
Yoon Kim, Alexander M Rush, Lei Yu, Adhiguna Kuncoro, Chris Dyer, and Gábor Melis. 2019 · 2019
Cited alongside, same era.
Multilingual constituency parsing with self-attention and pre-training
Nikita Kitaev, Steven Cao, and Dan Klein. 2019 · 2019
Cited alongside, same era.
PaLM: A hybrid parser and language model
Hao Peng, Roy Schwartz, and Noah A. Smith. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Ordered neurons: Integrating tree structures into recurrent neural networks
Yikang Shen, Shawn Tan, Alessandro Sordoni, and Aaron Courville. 2019 · 2019
Cited alongside, same era.
Peng Qian, Tahira Naseem, Roger Levy, and Ramón Fernandez Astudillo. 2021 · 2021
Later among the works it cites.
Self-attention networks can process bounded hierarchical languages
Shunyu Yao, Binghui Peng, Christos Papadimitriou, and Karthik Narasimhan. 2021 · 2021
Later among the works it cites.
Overcoming a theoretical limitation of self-attention
David Chiang and Peter Cholak. 2022 · 2022
Later among the works it cites.
Learning hierarchical structures with differentiable nondeterministic stacks
Brian DuSell and David Chiang. 2022 · 2022
Later among the works it cites.
Can transformers process recursive nested constructions, like humans?
Yair Lakretz, Théo Desbordes, Dieuwke Hupkes, and Stanislas Dehaene. 2022 · 2022
Later among the works it cites.
Transformer Grammars: Augmenting Transformer Language Models with Syntactic Inductive Biases at Scale
Laurent Sartran, Samuel Barrett, Adhiguna Kuncoro, Miloš Stanojević, Phil Blunsom, and Chris Dyer. 2022 · 2022
Later among the works it cites.
Can transformers learn to solve problems recursively?
Shizhuo Dylan Zhang, Curt Tigges, Stella Biderman, Maxim Raginsky, and Talia Ringer. 2023 · 2022
Later among the works it cites.
Neural networks and the chomsky hierarchy
Gregoire Deletang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein, Li Kevin Wenliang, Elliot Catt, Chris Cundy, Marcus Hutter, Shane Legg, Joel Veness, and Pedro A Ortega. 2023 · 2023
Closest in time.
Transformers learn shortcuts to automata
Bingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang. 2023 · 2023
Closest in time.
Characterizing intrinsic compositionality in transformers with tree projections
Shikhar Murty, Pratyusha Sharma, Jacob Andreas, and Christopher D Manning. 2023 · 2023
Closest in time.