Fetching the paper…
Reading the bibliography…
We focus on the recognition of Dyck-n ($\mathcal{D}_n$) languages with self-attention (SA) networks, which has been deemed to be a difficult task for these networks.
Memory-augmented recurrent neural networks can learn generalized dyck languages
Mirac Suzgun, Sebastian Gehrmann, Yonatan Belinkov, and Stuart M Shieber. 2019 · 1911
Earlier work this paper cites.
The algebraic theory of context-free languages
Noam Chomsky and Marcel P Schützenberger. 1959 · 1959
Earlier work this paper cites.
Finding structure in time
Jeffrey L Elman. 1990 · 1990
Earlier work this paper cites.
Learning context-free grammars: Capabilities and limitations of a recurrent neural network with an external stack memory
Sreerupa Das, C Lee Giles, and Guo-Zheng Sun. 1992 · 1992
Earlier work this paper cites.
On the computational power of neural nets
Hava T Siegelmann and Eduardo D Sontag. 1992 · 1992
Earlier work this paper cites.
A recurrent network that performs a context-sensitive prediction task
Mark Steijvers and Peter Grünwald. 1996 · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Lstm recurrent networks learn simple context-free and context-sensitive languages
Felix A Gers and E Schmidhuber. 2001 · 2001
Earlier work this paper cites.
Pretraining on non-linguistic structure as a tool for analyzing learning bias in language models
Isabel Papadimitriou and Dan Jurafsky. 2020 · 2004
Earlier work this paper cites.
What they do when in doubt: a study of inductive biases in seq2seq learners
Eugene Kharitonov and Rahma Chaabouni. 2020 · 2006
Cited alongside, same era.
A large annotated corpus for learning natural language inference
Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning. 2015 · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2015 · 2015
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Can recurrent neural networks learn nested recursion?
Jean-Philippe Bernardy. 2018 · 2018
Cited alongside, same era.
Context-free transductions with neural stacks
Yiding Hao, William Merrill, Dana Angluin, Robert Frank, Noah Amsel, Andrew Benz, and Simon Mendelsohn. 2018 · 2018
The importance of being recurrent for modeling hierarchical structure
Ke M Tran, Arianna Bisazza, and Christof Monz. 2018 · 2018
Later among the works it cites.
On the practical computational power of finite precision rnns for language recognition
Gail Weiss, Yoav Goldberg, and Eran Yahav. 2018 · 2018
Later among the works it cites.
On the turing completeness of modern neural network architectures
Jorge Pérez, Javier Marinković, and Pablo Barceló. 2019 · 2019
Later among the works it cites.
Ordered memory
Yikang Shen, Shawn Tan, Arian Hosseini, Zhouhan Lin, Alessandro Sordoni, and Aaron C Courville. 2019 · 2019
Later among the works it cites.
Learning the dyck language with attention-based seq2seq models
Xiang Yu, Ngoc Thang Vu, and Jonas Kuhn. 2019 · 2019
Later among the works it cites.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Listops: A diagnostic dataset for latent tree learning
Nikita Nangia and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
Evaluating the ability of LSTMs to learn context-free grammars
Luzi Sennhauser and Robert Berwick. 2018 · 2018
Cited alongside, same era.
Closing brackets with recurrent neural networks
Natalia Skachkova, Thomas Alexander Trost, and Dietrich Klakow. 2018 · 2018
Cited alongside, same era.
Closest in time.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. 2020 · 2020
Closest in time.
A formal hierarchy of rnn architectures
William Merrill, Gail Weiss, Yoav Goldberg, Roy Schwartz, Noah A Smith, and Eran Yahav. 2020 · 2020
Closest in time.