Fetching the paper…
Reading the bibliography…
Transformer architectures have been widely adopted in foundation models.
Memory-augmented recurrent neural networks can learn generalized dyck languages
M. Suzgun, S. Gehrmann, Y. Belinkov, and S. M. Shieber · 1911
Earlier work this paper cites.
Syntactic structures
N. Chomsky · 1957
Earlier work this paper cites.
The algebraic theory of context-free languages
N. Chomsky and M. P. Schützenberger · 1963
Earlier work this paper cites.
Finitary models of language users
G. A. Miller and N. Chomsky · 1963
Earlier work this paper cites.
Some complexity questions related to distributive computing (preliminary report)
A. C.-C. Yao · 1979
Earlier work this paper cites.
Extensions of lipschitz mappings into hilbert space
W. B. Johnson and J. Lindenstrauss · 1984
Earlier work this paper cites.
Communication Complexity
E. Kushilevitz and N. Nisan · 1996
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
An algebraic approach to communication complexity
J.-F. Raymond, P. Tesson, and D. Thérien · 1998
Earlier work this paper cites.
LSTM recurrent networks learn simple context-free and context-sensitive languages
F. A. Gers and E. Schmidhuber · 2001
Earlier work this paper cites.
A field guide to dynamical recurrent networks
J. F. Kolen and S. C. Kremer · 2001
Earlier work this paper cites.
Simple recurrent networks learn context-free and context-sensitive languages by counting
P. Rodriguez · 2001
Earlier work this paper cites.
Algorithmic derandomization via complexity theory
D. Sivakumar · 2002
Earlier work this paper cites.
Database-friendly random projections: Johnson-lindenstrauss with binary coins
D. Achlioptas · 2003
Earlier work this paper cites.
Constraints on multiple center-embedding of clauses
F. Karlsson · 2007
Earlier work this paper cites.
The one-way communication complexity of hamming distance
T. S. Jayram, R. Kumar, and D. Sivakumar · 2008
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
R. Pascanu, T. Mikolov, and Y. Bengio · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Structures, not strings: linguistics as part of the cognitive sciences
M. B. Everaert, M. A. Huybregts, N. Chomsky, R. C. Berwick, and J. J. Bolhuis · 2015
Earlier work this paper cites.
Using fast weights to attend to the recent past
J. Ba, G. E. Hinton, V. Mnih, J. Z. Leibo, and C. Ionescu · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Closing brackets with recurrent neural networks
N. Skachkova, T. A. Trost, and D. Klakow · 2018
Earlier work this paper cites.
On the distribution of deep clausal embeddings: A large cross-linguistic study
D. Blasi, R. Cotterell, L. Wolf-Sonkin, S. Stoll, B. Bickel, and M. Baroni · 2019
Earlier work this paper cites.
On the computational power of rnns
S. A. Korsky and R. C. Berwick · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al · 2019
Cited alongside, same era.
On the turing completeness of modern neural network architectures
J. Pérez, J. Marinković, and P. Barceló · 2019
Cited alongside, same era.
Learning the dyck language with attention-based seq2seq models
X. Yu, N. T. Vu, and J. Kuhn · 2019
Cited alongside, same era.
On the Ability and Limitations of Transformers to Recognize Formal Languages
S. Bhattamishra, K. Ahuja, and N. Goyal · 2020
Cited alongside, same era.
On the practical ability of recurrent neural networks to recognize hierarchical languages
Neural networks and the chomsky hierarchy
G. Deletang, A. Ruoss, J. Grau-Moya, T. Genewein, L. K. Wenliang, E. Catt, C. Cundy, M. Hutter, S. Legg, J. Veness, and P. A. Ortega · 2023
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
A. Gu and T. Dao · 2023
Later among the works it cites.
Transformers learn shortcuts to automata
B. Liu, J. T. Ash, S. Goel, A. Krishnamurthy, and C. Zhang · 2023
Later among the works it cites.
The parallelism tradeoff: Limitations of log-precision transformers
W. Merrill and A. Sabharwal · 2023
Later among the works it cites.
Resurrecting recurrent neural networks for long sequences
A. Orvieto, S. L. Smith, A. Gu, A. Fernando, C. Gulcehre, R. Pascanu, and S. De · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Bhattamishra, K. Ahuja, and N. Goyal · 2020
Cited alongside, same era.
On the computational power of transformers and its implications in sequence modeling
S. Bhattamishra, A. Patel, and N. Goyal · 2020
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Cited alongside, same era.
How can self-attention networks recognize Dyck-n languages?
J. Ebrahimi, D. Gelda, and W. Zhang · 2020
Cited alongside, same era.
Theoretical limitations of self-attention in neural sequence models
M. Hahn · 2020
Cited alongside, same era.
RNNs can generate bounded hierarchical languages with optimal memory
J. Hewitt, M. Hahn, S. Ganguli, P. Liang, and C. D. Manning · 2020
Cited alongside, same era.
Transformers are rnns: Fast autoregressive transformers with linear attention
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret · 2020
Cited alongside, same era.
B. Peng, E. Alcaide, Q. Anthony, A. Albalak, S. Arcadinho, H. Cao, X. Cheng, M. Chung, M. Grella, K. K. GV, et al · 2023
Later among the works it cites.
Representational strengths and limitations of transformers
C. Sanford, D. Hsu, and M. Telgarsky · 2023
Later among the works it cites.
Retentive network: A successor to transformer for large language models
Y. Sun, L. Dong, S. Huang, S. Ma, Y. Xia, J. Xue, J. Wang, and F. Wei · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
J. Von Oswald, E. Niklasson, E. Randazzo, J. Sacramento, A. Mordvintsev, A. Zhmoginov, and M. Vladymyrov · 2023
Later among the works it cites.
What algorithms can transformers learn? a study in length generalization
H. Zhou, A. Bradley, E. Littwin, N. Razin, O. Saremi, J. Susskind, S. Bengio, and P. Nakkiran · 2023
Later among the works it cites.
Zoology: Measuring and improving recall in efficient language models
S. Arora, S. Eyuboglu, A. Timalsina, I. Johnson, M. Poli, J. Zou, A. Rudra, and C. Re · 2024
Closest in time.
Transformers as statisticians: Provable in-context learning with in-context algorithm selection
Y. Bai, F. Chen, H. Wang, C. Xiong, and S. Mei · 2024
Closest in time.
Understanding in-context learning in transformers and LLMs by learning to learn discrete functions
S. Bhattamishra, A. Patel, P. Blunsom, and V. Kanade · 2024
Closest in time.
Griffin: Mixing gated linear recurrences with local attention for efficient language models
S. De, S. L. Smith, A. Fernando, A. Botev, G. Cristian-Muraru, A. Gu, R. Haroun, L. Berrada, Y. Chen, S. Srinivasan, et al · 2024
Closest in time.
Repeat after me: Transformers are better than state space models at copying
S. Jelassi, D. Brandfonbrener, S. M. Kakade, and E. Malach · 2024
Closest in time.
Tracr: Compiled transformers as a laboratory for interpretability
D. Lindner, J. Kramár, S. Farquhar, M. Rahtz, T. McGrath, and V. Mikulik · 2024
Closest in time.
Exposing attention glitches with flip-flop language modeling
B. Liu, J. Ash, S. Goel, A. Krishnamurthy, and C. Zhang · 2024
Closest in time.
A logic for expressing log-precision transformers
W. Merrill and A. Sabharwal · 2024
Closest in time.
On limitations of the transformer architecture, 2024
B. Peng, S. Narayanan, and C. Papadimitriou · 2024
Closest in time.
Transformers, parallel computation, and logarithmic depth
C. Sanford, D. Hsu, and M. Telgarsky · 2024
Closest in time.
What formal languages can transformers express? a survey
L. Strobl, W. Merrill, G. Weiss, D. Chiang, and D. Angluin · 2024
Closest in time.
Transformers are uninterpretable with myopic methods: a case study with bounded dyck grammars
K. Wen, Y. Li, B. Liu, and A. Risteski · 2024
Closest in time.